María Engracia Gómez

dblp:18/79 · also María Engracia Gómez Requena · DBLP profile ↗
← Back
78ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0003-1466-4118ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 65 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author
YearPublicationVenuePosition
2026 SYNPA: understanding the impact of the interference modeling on thread-to-core allocation policies for SMT ARM processors
abstract
Modern high-performance servers increasingly rely on Simultaneous Multithreading (SMT) processors to enhance throughput with minimal area overhead. However, SMT architectures introduce inter-application interference, often resulting in degraded performance for individual applications. To address this issue, interference-aware thread-to-core (T2C) allocation policies are essential. This paper explores the design and implementation of such policies using real performance counters on ARM processors. We introduce the Instructions and Stalls Cycles (ISC) stack—a simple yet effective model for characterizing application behavior and identifying synergistic thread pairings. Building on our previous work, SYNPA, we improve the accuracy of the model by accounting for horizontal waste (that is, unused dispatch slots) and proposing methods to address limitations in ARM’s Performance Monitoring Unit (PMU), which prevent complete attribution of processor cycles. These enhancements result in a family of SYNPA schedulers, each based on a different ISC stack variant. Detailed discussions are provided on the pros and cons that researchers typically face when building a performance stack on commercial processors. These analyses are intended to assist researchers in their work.
Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
J. Supercomput.4
2026 WAPA: A Microarchitecture- and Workload-Agnostic Universal SMT Scheduler
Marta Navarro 0001, Vicent Pallardó-Julià, Lucia Pons, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
IEEE Trans. Parallel Distributed Syst.5
2025 WAPA: A Workload-Agnostic CPI-Based Thread-to-Core Allocation Policy
Marta Navarro 0001, Vicent Pallardó-Julià, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
Euro-Par (1)4
2025 Power, energy, and performance analysis of single- and multi-threaded applications in the ARM ThunderX2
abstract
Energy efficiency has been a major concern in data centers, and the problem is exacerbated as its size continues to rise. However, the lack of tools to measure and handle this energy at a fine granularity (e.g., processor core or last-level cache) has translated into slow research advances in this topic. Understanding where (i.e., which components) and when (the point in time) energy consumption translates into minor performance improvements is of paramount importance to design any energy-aware scheduler. This paper characterizes the relationship between energy consumption and performance in a 28-core ARM ThunderX2 processor for both single-threaded and multi-threaded applications. This paper shows that single-threaded applications with high CPU activity maintain their performance in spite of the inter-application interference at shared resources, but this comes at the expense of higher power consumption. Conversely, applications that heavily utilize the L3 cache and memory consume less power but suffer significant performance degradation as interference levels rise. In contrast, multi-threaded applications show two distinct behaviors. On the one hand, some of them experience significant performance gains when they execute in a higher number of cores with more threads, which outweighs the increase in power consumption, leading to high energy efficiency.
Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
J. Parallel Distributed Comput.3
2024 Optimization of One-to-Many Communication Primitives for Dragonfly Topologies
abstract
Collective communication primitives (CCPs), such as multicast and broadcast, are essential for many parallel and distributed applications. In response, this study compares a topology-oblivious algorithm underlying the implementation of CCPs in standard instances of MPI with two topology-aware implementations, based on the LLF and GLF algorithms, and an ideal hardware-assisted approach. By using real scientific applications, instead of synthetic traffic, our study reveals workload-dependent performance variations among CCP implementations; and highlights the importance of CCP algorithm selection in optimizing application performance in supercomputing environments.
Jose Duro, Adrián Castelló 0001, María Engracia Gómez, Julio Sahuquillo, Enrique S. Quintana-Ortí
ICPADS3
2024 SYNPA: SMT Performance Analysis and Allocation of Threads to Cores in ARM Processors
abstract
Simultaneous multithreading processors improve throughput over single-threaded processors thanks to sharing internal core resources among instructions from distinct threads. However, resource sharing introduces inter-thread interference within the core, which has a negative impact on individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the interference effects depends on the competing co-runners sharing the core. Thus, it can be mitigated by applying a thread-to-core allocation policy that smartly selects applications to be run in the same core to minimize their interference.This paper presents SYNPA, a simple approach that dynamically allocates threads to cores in an SMT processor based on their run-time dynamic behavior. The approach uses a regression model to select synergistic pairs to mitigate intra-core interference. The main novelty of SYNPA is that it uses just three variables collected from the performance counters available in current ARM processors at the dispatch stage. Experimental results show that SYNPA outperforms the default Linux scheduler by around 36%, on average, in terms of turnaround time in 8-application workloads combining frontend-bound and backend-bound benchmarks.
Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
IPDPS4
2024 Characterizing Power and Performance Interference Scalability in the 28-core ARM ThunderX2
abstract
Nowadays, energy efficiency is a major concern in any type of processor-based device, ranging from processor servers to supercomputers, including mobile battery-fed devices. In recent years, systems based on the ARM architecture, tra-ditionally better suited for mobile and embedded systems, have increased their market share in segments commonly occupied by x86 processors. This growth is due, to some extent, to the excellent energy efficiency shown by high-performance ARM processors. Designing software and hardware energy-efficient systems requires a sound knowledge of the relationship among three main axes: component activity, power, and inter-application interference at the shared resources. This problem aggravates in many-core processors, which are ubiquitous in high-performance servers. This paper characterizes the aforementioned axes in a 28-core ARM Thunder X2 processor. Experimental results show that the performance of highly-scalable single-threaded applications is sustained regardless of the number of applications and interference introduced at the shared resources at the cost of increasing power. In contrast, low-scalable applications require much less power and experience a huge performance degradation as their number increases.
Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
PDP3
2024 A modular approach to build a hardware testbed for cloud resource management research
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo
J. Supercomput.4
2023 Thread-to-Core Allocation in ARM Processors Building Synergistic Pairs
abstract
Simultaneous multithreading (SMT) processors can present significant throughput improvements over single-threaded (ST) processors thanks to sharing internal core resources among instructions executing from multiple threads. However, resource sharing introduces inter-thread interference within the core, which negatively impacts individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the intra-core interference on performance depends on the applications co-running in the same core. A thread scheduler can help reduce this effect by smartly selecting the pairs of applications that should run on each SMT core. This paper presents SYNPA, a simple approach that dynamically allocates threads to SMT cores based on their run-time dynamic behavior. SYNPA uses a regression model to select synergistic pairs to mitigate intra-core interference. Results show that SYNPA outperforms the default Linux scheduler by around 35%, on average, in terms of turnaround time when running 8-application workloads combining frontend-bound and backend-bound applications.
Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
PACT4
2023 Stratus: A Hardware/Software Infrastructure for Controlled Cloud Research
abstract
Cloud systems deploy a wide variety of shared resources and host a large number of tenant applications. To perform cloud research, a small experimental platform is commonly used, which hides the huge system complexity and provides flexibility. Despite being simpler, this platform should include the main cloud system components (hardware and software) to provide representative results. A wide set of platforms have spread in recent years; however, most of them only include a major cloud component or lack the deployment of virtual machines (VMs) to provide isolation. This paper presents Stratus, an experimental platform that is currently being used to carry out cloud research. To the best of our knowledge, Stratus is the only platform that jointly provides three main features: uses VMs to isolate tenant applications, deploys the three types of cloud nodes (server, client, and storage), and manages all main shared system resources (CPUs, LLC space, memory, network, and disk bandwidth). Moreover, Stratus implements a software manager to ease the research and aid the design of QoS-aware policies. The manager integrates three main functionalities: management and control of the execution of VMs and running applications, monitoring of hardware performance counters and system resource utilization, and partitioning of the main shared system resources by using technologies available in commercial processors.
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Chaoyi Huang, Julio Sahuquillo
PDP4
2023 Cloud White: Detecting and Estimating QoS Degradation of Latency-Critical Workloads in the Public Cloud
abstract
The increasing popularity of cloud computing has forced cloud providers to build economies of scale to meet the growing demand. Nowadays, data-centers include thousands of physical machines, each hosting many virtual machines (VMs), which share the main system resources, causing interference that can significantly impact on performance. Frequently, these data-centers run latency-critical workloads, whose performance is determined by tail latency, which is very sensitive to the interference of co-running workloads. To prevent QoS violations, cloud providers adopt overprovisioning strategies but they reduce the server utilization and increase the costs. A mechanism that accurately estimates performance degradation dynamically in a production system would allow cloud providers to improve the servers’ utilization. In this work we propose Cloud White, an approach that is able to detect the inter-VM interference in scenarios with multiple co-located latency-critical VMs and estimate the performance degradation using multi-variable regression models. Unlike previous proposals, Cloud White is built taking into account the limitations of a public cloud production system. Experimental results show that Cloud White is able to estimate performance degradation with a small overall prediction error of 5%.
Lucia Pons, Josué Feliu, Julio Sahuquillo, María Engracia Gómez, Salvador Petit, Julio Pons, Chaoyi Huang
Future Gener. Comput. Syst.4
2022 RED-SEA: Network Solution for Exascale Architectures
abstract
In order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions.
Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile
DSD65
2022 A Neural Network to Estimate Isolated Performance from Multi-Program Execution
abstract
When multiple applications are running on a platform with shared resources like multicore CPUs, the behaviour of the running application can be altered by the co-runners. In this case, the system resources need to be managed (e.g. by repartitioning the cache space, re-schedule applications in distinct cores, modifying the prefetcher configuration, etc.) to reduce the inter-application interference in order to minimize the performance losses over isolated execution. In this context, a main challenge in different computing scenarios like the public cloud or soft real-time systems is knowing the performance impact of a given management action on each application with respect to its isolated execution. With this aim, in this work we present a neural network-based approach that estimates the performance an application would have had in isolation from multi-program executions. Experimental results show that the proposal dynamically adapts to changes in application behavior. On average, the predicted performance presents an error deviation by 11.7% and 2.3% for MAPE and MSE respectively.
Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
PDP4
2022 Effect of Hyper-Threading in Latency-Critical Multithreaded Cloud Applications and Utilization Analysis of the Major System Resources
abstract
Multithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main system resources is a major concern for cloud providers in order to devise virtualization strategies that improve the system efficiency. With this aim, this paper first characterizes the impact of QPS on tail latency, analyzing different scenarios varying the number of threads and the thread-to-core allocation (single-task and multi-task execution) policy. The characterization study reveals that the performance of some applications does not scale with the number of threads, and the performance of some others is insensitive to the Hyper-Threading technology, so they can be allocated in less physical cores and improve system utilization. Identifying these applications, however, at run-time is challenging. Despite identifying these applications at run-time is challenging, this paper shows that they can be successfully detected at run-time by analyzing the utilization trend of the major system resources. In addition to CPU, we have also studied how assigning the share of each application of other major shared system resources impacts on performance. We outline considerations cloud providers should take into account to improve performance and resource utilization.
Lucia Pons, Josué Feliu, José Puche, Chaoyi Huang, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo
Future Gener. Comput. Syst.7
2022 DeepP: Deep Learning Multi-Program Prefetch Configuration for the IBM POWER 8
abstract
Current multi-core processors implement sophisticated hardware prefetchers, that can be configured by application (PID), to improve the system performance. When running multiple applications, each application can present different prefetch requirements, hence different configurations can be used. Setting the optimal prefetch configuration for each application is a complex task since it does not only depend on the application characteristics but also on the interference at the shared memory resources (e.g., memory bandwidth). In his paper, we proposeDeepP, a deep learning approach for the IBM POWER8 that identifies at run-time the best prefetch configuration for each application in a workload. To this end, the neural network predicts the performance of each application under the studied prefetch configurations by using a set of performance events. The prediction accuracy of the network is improved thanks to a dynamic training methodology that allows learning the impact of dynamic changes of the prefetch configuration on performance. At run-time, the devised network infers the best prefetch configuration for each application and adjusts it dynamically. Experimental results show that the proposed approach improves performance, on average, by 5.8%, 6.7%, and 15.8% compared to the default prefetch configuration across different 6-, 8-, and 10-application workloads, respectively.
Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
IEEE Trans. Computers4
2020 Impact of the Array Shape and Memory Bandwidth on the Execution Time of CNN Systolic Arrays
abstract
The use of Convolutional Neural Networks (CNN) has experienced a huge rise over the last recent years and its popularity has increased exponentially, mainly due to its application both for image recognition and certain applications related to artificial intelligence. The new applications of CNN request computing demands that are difficult to address by conventional processors.As a consequence, accelerators -both prototypes and commercial products- focusing on CNN computation have been proposed. Among these accelerators, those based on systolic arrays have acquired a special relevance; some examples are the Google's TPU and Eyeriss.Current research has focused on regular squared systolic arrays and most existing work assumes that there is enough memory bandwidth to feed the systolic array with input data. In this paper we explore the design of non-squared systolic arrays and address the impact of the memory bandwidth from a performance perspective.This work makes two main contributions. First, we found that some workloads with non-squared arrays achieve similar performance to systolic arrays twice as large, which can translate in area and/or energy benefits.Second, we present a performance comparison varying the main memory bandwidth for current DRAM devices. The analysis reveals that main memory bandwidth has a great impact on performance and that the decision of which technology use is key for the system performance. For the 64x64 array size it is necessary to use HBM2 memory to avoid the slowdown that would introduce cheaper technologies (e.g. DDR5 and DDR4).
Eduardo Yago, Pau Castelló, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
DSD4
2020 An efficient cache flat storage organization for multithreaded workloads for low power processors
José Puche, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
Future Gener. Comput. Syst.3
2020 Bandwidth-Aware Dynamic Prefetch Configuration for IBM POWER8
abstract
Advanced hardware prefetch engines are being integrated in current high-performance processors. Prefetching can boost the performance of most applications, however, the induced bandwidth consumption can lead the system to a high contention for main memory bandwidth, which is a scarce resource in current multicores. In such a case, the system performance can be severely damaged. This article characterizes the applications’ behavior in an IBM POWER8 machine, which presents many prefetch settings, varying the bandwidth contention. The study reveals that the best prefetch setting for each application depends on the main memory bandwidth availability, that is, it depends on the co-running applications. Based on this study, we propose Bandwidth-Aware Prefetch Configuration (BAPC) a scalable adaptive prefetching algorithm that improves the performance of multi-program workloads. BAPC increases the performance of the applications in a 12, 15, and 16 percent of 6-, 8-, and 10-application workloads over the IBM POWER8 default configuration. In addition, BAPC reduces bandwidth consumption in 39, 42, and 45 percent, respectively.
Carlos Navarro, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
IEEE Trans. Parallel Distributed Syst.4
2019 Modeling and analysis of the performance of exascale photonic networks
abstract
Summary Photonics technology has become a promising and viable alternative for both on‐chip and off‐chip interconnection networks of future Exascale systems. Nevertheless, this technology is not mature enough yet in this context, so research efforts focusing on photonic networks are still required to achieve realistic suitable network implementations. In this regard, system‐level photonic network simulators can help guide designers to assess the multiple design choices. Most current research is done on electrical network simulators, whose components work widely different from photonics components. In this work, we summarize and compare the working behavior of both technologies which includes the use of optical routers, wavelength‐division multiplexing and circuit switching among others. After implementing them into a well‐known simulation framework, an extensive simulation study has been carried out using realistic photonic network configurations with synthetic and realistic traffic. Experimental results show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. Our study also reveals that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters.
Jose Duro, Jose Antonio Pascual, Salvador Petit, Julio Sahuquillo, María Engracia Gómez
Concurr. Comput. Pract. Exp.5
2019 FOS: a low-power cache organization for multicores
José Puche, Salvador Petit, Julio Sahuquillo, María Engracia Gómez
J. Supercomput.4
2018 Efficient selective multicore prefetching under limited memory bandwidth
Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena
J. Parallel Distributed Comput.3
2018 TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction
abstract
Discerning the private or shared condition of the data accessed by the applications is an increasingly decisive approach to achieving efficiency and scalability in multiand many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor. This paper introduces TokenTLB, a TLB-based page classification approach based on exchange and count of tokens. Token counting on TLBs is a natural and efficient way for classifying memory pages, and it does not require the use of complex and undesirable persistent requests or arbitration. In addition, classification is extended with Cooperative Usage Predictor (CUP), a token-based system-wide page usage predictor retrieved through TLB cooperation, in order to perform a classification unaffected by TLB size. Through cycle-accurate simulation we observed that TokenTLB spends 43.6 percent of cycles as private per page on average, and CUP further increases the time spent as private by 22.0 percent. CUP avoids 4 out of 5 TLB invalidations when compared to state-of-the-art predictors, thus proving far better prediction accuracy and making usage prediction an attractive mechanism for the first time.
Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez
IEEE Trans. Parallel Distributed Syst.4
2017 Application Clustering Policies to Address System Fairness with Intel's Cache Allocation Technology
abstract
Achieving system fairness is a major design concern in current multicore processors. Unfairness arises due to contention in the shared resources of the system, such as the LLC and main memory. To address this problem, many research works have proposed novel cache partitioning policies aimed at addressing system fairness without harming performance. Unfortunately, existing proposals targeting fairness require extra hardware which makes them impractical in commercial processors.Recent Intel Xeon processors feature Cache Allocation Technology (CAT), a hardware cache partitioning mechanism that can be controlled from userspace software and that allows to create partitions in the LLC and assign different groups of applications to them.In this paper we propose a family of clustering-based cache partitioning policies to address fairness in systems that feature Intel's CAT. The proposal acts at two levels: applications showing similar amount of core stalls due to LLC accesses are first grouped into clusters, after which each cluster is given a number of ways using a simple mathematical model. To the best of our knowledge, this is the first attempt to address system fairness using the cache partitioning hardware in a real product. Results show that our best performing policy reduces system unfairness by up to 80% (39% on average) for 8-application workloads and by up to 45% (25% on average) for 12-application workloads compared to a non-partitioning approach.
Vicent Selfa, Julio Sahuquillo, Lieven Eeckhout, Salvador Petit, María Engracia Gómez
PACT5
2017 A fault-tolerant routing strategy for k-ary n-direct s-indirect topologies based on intermediate nodes
abstract
Summary Exascale computing systems are being built with thousands of nodes. The high number of components of these systems significantly increases the probability of failure. A key component for them is the interconnection network. If failures occur in the interconnection network, they may isolate a large fraction of the machine. For this reason, an efficient fault‐tolerant mechanism is needed to keep the system interconnected, even in the presence of faults. A recently proposed topology for these large systems is the hybridk‐aryn‐directs‐indirect family that provides optimal performance and connectivity at a reduced hardware cost. This paper presents a fault‐tolerant routing methodology for thek‐aryn‐directs‐indirect topology that degrades performance gracefully in presence of faults and tolerates a large number of faults without disabling any healthy computing node. In order to tolerate network failures, the methodology uses a simple mechanism. For any source‐destination pair, if necessary, packets are forwarded to the destination node through a set of intermediate nodes (without being ejected from the network) with the aim of circumventing faults. The evaluation results shows that the proposed methodology tolerates a large number of faults. For instance, it is able to tolerate more than 99.5% of fault combinations when there are 10 faults in a 3‐D network with 1000 nodes using only 1 intermediate node and more than 99.98% if 2 intermediate nodes are used. Furthermore, the methodology offers a gracious performance degradation. As an example, performance degrades only by 1% for a 2‐D network with 1024 nodes and 1% faulty links.
Roberto Peñaranda, María Engracia Gómez, Pedro López 0001, Ernst Gunnar Gran, Tor Skeie
Concurr. Comput. Pract. Exp.2
2017 A research-oriented course on Advanced Multicore Architecture: Contents and active learning methodologies
Salvador Petit, Julio Sahuquillo, María Engracia Gómez, Vicent Selfa
J. Parallel Distributed Comput.3
2017 The Tag Filter Architecture: An energy-efficient cache and directory design
Joan J. Valls, Alberto Ros 0001, María Engracia Gómez, Julio Sahuquillo
J. Parallel Distributed Comput.3
2017 XOR-based HoL-blocking reduction routing mechanisms for direct networks
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001
Parallel Comput.3
2017 TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs
abstract
Recent proposals are based on classifying memory accesses into private or shared in order to process private accesses more efficiently and reduce coherence overhead. The classification mechanisms previously proposed are either not able to adapt to the dynamic sharing behavior of the applications or require frequent broadcast messages. Additionally, most of these classification approaches assume single-level translation lookaside buffers (TLBs). However, deeper and more efficient TLB hierarchies, such as the ones implemented in current commodity processors, have not been appropriately explored. This paper analyzes accurate classification mechanisms in multilevel TLB hierarchies. In particular, we propose an efficient data classification strategy for systems with distributed shared last-level TLBs. Our approach classifies data accounting for temporal private accesses and constrains TLB-related traffic by issuing unicast messages on first-level TLB misses. When our classification is employed to deactivate coherence for private data in directory-based protocols, it improves the directory efficiency and, consequently, reduces coherence traffic to merely 53.0 percent, on average. Additionally, it avoids some of the overheads of previous classification approaches for purely private TLBs, improving average execution time by nearly 9 percent for large-scale systems.
Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
IEEE Trans. Parallel Distributed Syst.3
2017 A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared Caches
abstract
Shared caches have become the common design choice in the vast majority of modern multi-core and many-core processors, since cache sharing improves throughput for a given silicon area. Sharing the cache, however, has a downside: the requests from multiple applications compete among them for cache resources, so the execution time of each application increases over isolated execution. The degree in which the performance of each application is affected by the interference becomes unpredictable yielding the system to unfairness situations. This paper proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead hardware-based cache partitioning approach that addresses system fairness. FPCP reduces the interference by allocating to each application a cache partition and adjusting the partition sizes at runtime. To adjust partitions, our approach estimates during multicore execution the time each application would have taken in isolation, which is challenging. The proposed approach has two main differences over existing approaches. First, FPCP distributes cache ways incrementally, which makes the proposal less prone to estimation errors. Second, the proposed algorithm is much less costly than the state-of-the-art ASM-Cache approach. Experimental results show that, compared to ASM-Cache, FPCP reduces unfairness by 48 percent in four-application workloads and by 28 percent in eight-application workloads, without harming the performance.
Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez
IEEE Trans. Parallel Distributed Syst.4
2016 Student Research Poster: A Low Complexity Cache Sharing Mechanism to Address System Fairness
abstract
Shared caches have become, de facto, the common design choice in current multi-cores, ranging from embedded devices to high-performance processors. In these systems, requests from multiple applications compete for the cache resources, degrading to different extents their progress, quantified as the performance of individual applications compared to isolated execution. The difference between the progresses of the running applications yields the system to unpredictable behavior and causes a fairness problem. This problem can be addressed by carefully partitioning cache resources among the contending applications, but to be effective, a partitioning approach needs to estimate per-application progress. This work proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead cache partitioning approach which addresses fairness by distributing cache resources among applications depending on their progress. To estimate progress, we have implemented two state-of-the-art performance models, ASM and PTCA, which estimate, at runtime, the performance a given application would have if executed in isolation.
Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez
PACT4
2016 A Directory Cache with Dynamic Private-Shared Partitioning
abstract
As the core counts increase in each chip multiprocessor generation, coherence protocols should improve scalability inperformance, area, and energy consumption to meet the demandsof larger core counts. Directory-based protocols constitute themost scalable alternative. A conventional directory, however, suffers from an inefficient use of storage and energy. First, thelarge, non-scalable, sharer vectors consume unnecessary area andleakage, especially considering that most of the blocks trackedin a directory are cached by a single core. Second, althoughincreasing directory size and associativity could boost systemperformance, it would come at expenses of energy consumption. This paper proposes the Dynamic Way Partitioning (DWP) Directory, a directory structure that exploits three main workloadcharacteristics to achieve area and energy reductions. First, it iswidely known that even in parallel workloads most of the accessedcache blocks are private. Second, most directory accesses targetthe small number of shared blocks. Third, the shared/privateratio of entries in the directory varies across applications andacross different execution phases within the applications. To takeadvantage of these three characteristics, DWP-Directory reducesthe number of ways with storage for shared blocks and it allowsthis storage to be powered off or on at run-time according to thedynamic requirements of the applications. DWP-Directory is compared to a conventional directory cachewith different associativity degrees and with two state-of-the-artschemes: PS-Directory and Hybrid Representation. Experimentalresults for 32-core CMPs show that DWP-Directory achievesthe best of both worlds: similar performance as a traditionaldirectory with high associativity, and similar area as recentstate-of-the-art schemes. In addition, DWP-Directory reducesstatic and dynamic power consumption by 38.0% and 67.4%,respectively compared to conventional sparse directories.
Joan J. Valls, María Engracia Gómez, Alberto Ros 0001, Julio Sahuquillo
HiPC2
2016 TokenTLB: A Token-Based Page Classification Approach
abstract
Classifying memory accesses into private or shared data has become a fundamental approach to achieving efficiency and scalability in multi- and many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor.
Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez, José Duato
ICS4
2016 A Simple Activation/Deactivation Prefetching Scheme for Chip Multiprocessors
abstract
Prefetching significantly reduces the memory latencies of a wide range of applications and thus increases the system performance. However, as a speculative technique, prefetching may also noticeably increase the number of memory accesses, which in turns may negatively impact on the main memory bandwidth consumption, performance, and power. Main memory bandwidth consumption is a critical resource especially in the context of current multicore processors since memory requests from all the cores, both prefetch and demand requests, compete among them in the access to the DRAM banks. Consequently, demand requests may be delayed hurting the system performance. This work proposes the Activation/Deactivation Policies (ADP) scheme for hardware prefetchers in multicore processors. This scheme relies on activation policies that turn on the prefetcher on a given core when it is expected that prefetches will improve the performance, and turn off the prefetcher of that core when it is foreseen that performance will be scarcely improved or not improved at all. The proposed mechanism effectively reduces the memory bandwidth requirements of some cores with respect to a typical always prefetching mechanism, so making available extra bandwidth to the co-runners. Results in a four-core processor show that ADP prefetching achieves similar performance ±2.5% as always prefetching, while significantly reducing the memory bandwidth consumed by use-less prefetches. Moreover, in some applications this reduction is as much as 50%. ADP prefetching is applicable to stream-based prefetchers, global-history-buffer delta correlation prefetchers, and PC-based stride prefetchers.
Vicent Selfa, Crispín Gómez Requena, María Engracia Gómez, Julio Sahuquillo
PDP3
2016 The k-ary n-direct s-indirect family of topologies for large-scale interconnection networks
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
J. Supercomput.3
2016 Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors
abstract
Most of the data referenced by sequential and parallel applications running in current chip multiprocessors are referenced by a single thread, i.e., private. Recent proposals leverage this observation to improve many aspects of chip multiprocessors, such as reducing coherence overhead or the access latency to distributed caches. The effectiveness of those proposals depends to a large extent on the amount of detected private data. However, the mechanisms proposed so far do not consider neither thread migration nor the private use of data within different application phases. As a result, a considerable amount of private data is not detected. In order to increase the detection of private data, we propose a TLB-based mechanism that is able to account for both thread migration and application phases. Simulation results show that the average number of pages detected as private significantly increases from 43 percent in previous proposals up to 79 percent in ours while keeping a reasonable TLB miss rate. Furthermore, when our proposal is used to deactivate the coherence for private data in a directory protocol, it improves execution time by 13.5 percent, on average, with respect to previous techniques.
Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
IEEE Trans. Parallel Distributed Syst.3
2016 A Family of Fault-Tolerant Efficient Indirect Topologies
abstract
On the one hand, performance and fault-tolerance of interconnection networks are key design issues for high performance computing (HPC) systems. On the other hand, cost should be also considered. Indirect topologies are often chosen in the design of HPC systems. Among them, the most commonly used topology is the fat-tree. In this work, we focus on getting the maximum benefits from the network resources by designing a simple indirect topology with very good performance and fault-tolerance properties, while keeping the hardware cost as low as possible. To do that, we propose some extensions to the fat-tree topology to take full advantage of the hardware resources consumed by the topology. In particular, we propose three new topologies with different properties in terms of cost, performance and fault-tolerance. All of them are able to achieve a similar or better performance results than the fat-tree, providing also a good level of fault-tolerance and, contrary to most of the available topologies, these proposals are able to tolerate also faults in the links that connect to end nodes.
Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
IEEE Trans. Parallel Distributed Syst.3
2015 Row Tables: Design Choices to Exploit Bank Locality in Multiprogram Workloads
abstract
Main memory is a major performance bottleneck in current chip multiprocessors. Current DRAM banks latch the last accessed row in an internal buffer, namely row buffer (RB), which allows fast subsequent accesses to that row. This throughput-oriented approach was originally designed for single-thread processors and pursues to take advantage of the spatial locality that individual applications exhibit. This paper proposes row tables, a pool of row buffers shared among threads. Depending on the needs of each thread, row buffers are dynamically allocated to threads. Two design approaches are devised differing on the table location, and referred to as BRT (Bank Row Table) and CRT (Controller Row Table), which place the table at the bank, as traditionally done in existing modules, and at the memory controller side, respectively. CRT performs better than BRT in high RB locality applications (or mixes) but performs worse in poor RB locality applications since the increase in transfer times is not later amortized. A variant of CRT referred to as CRT1/xhas been devised to reduce this performance penalty. Results for a 4-core system show that, on average, BRT and CRT1/xmechanisms save energy by 23% and 7%-16% (depending on the X value) and improve IPC by 10% and 9%-14%, respectively.
Paula Navarro, Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena
PDP4
2015 XORAdap: A HoL-Blocking Aware Adaptive Routing Algorithm
abstract
Routing is a key parameter in the design of the interconnection network of large parallel computers. Depending on the number of routing options available for each packet, routing algorithms can be deterministic (one available path) or adaptive (several ones). Adaptive routing usually outperforms deterministic routing but it also may increase the Head-of-Line blocking effect. Usually, adaptive routing uses virtual channels to provide routing flexibility and to guarantee deadlock freedom. On the other hand, deterministic routing is simpler and therefore it has lower routing delay. In this paper, we take the challenge of developing new routing algorithms for direct topologies that exploit virtual channels in an efficient way combining the good properties of both routing algorithms types: flexibility and reduced HoL blocking. To do that, this paper proposes several hybrid (combination of adaptivity and determinism) simple mechanisms to perform an efficient distribution of packets among virtual channels based on their destination. The resulting routing mechanisms are able to adapt to the different traffic patterns to obtain the best performance while keeping the simplicity of routing.
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001
PDP3
2015 Methodologies and Performance Metrics to Evaluate Multiprogram Workloads
abstract
Multicore processors are dominating the microprocessor market and most research work has moved to this kind of processors. Multicore research methods are still immature and evolving from the single-threaded processor ounterparts. Three main research issues must be faced when evaluating performance and energy in multicores. First, multiple simulation methodologies are being applied to evaluate these systems, without being an agreement about which to use. Second, due to the nature of multiprogram workloads new performance metrics are required, different from those used in single-thread processors. Many metrics have been defined and distinct metrics are used across the published works. Finally, multicore processors are really complex systems which require from sophisticated and complementary (e.g. energy and performance) simulators. This paper pursues to help researchers face the three mentioned research issues. For this purpose, we compare these issues across 28 papers published in 2013 in top computer architecture conferences. Both analytical examples and experimental results are presented with the aim of providing some insights in multicore research.
Vicent Selfa, Julio Sahuquillo, Crispín Gómez Requena, María Engracia Gómez
PDP4
2015 The Tag Filter Cache: An Energy-Efficient Approach
abstract
Power consumption in current high-performance chip multiprocessors (CMPs) has become a major design concern. The current trend of increasing the core count aggravates this problem. On-chip caches consume a significant fraction of the total power budget. Most of the proposed techniques to reduce the energy consumption of these memory structures are at the cost of performance, which may become unacceptable for high-performance CMPs. On-chip caches in multi-core systems are usually deployed with a high associativity degree in order to enhance performance. Even first-level caches are currently implemented with eight ways. The concurrent access to all the ways in the cache set is costly in terms of energy. In this paper we propose an energy-efficient cache design, namely the Tag Filter Cache (TF-Cache) architecture, that filters some of the set ways during cache accesses, allowing to access only a subset of them without hurting the performance. Our cache for each way stores the lowest order tag bits in an auxiliary bit array and these bits are used to filter the ways that do not match those bits in the searched block tag. Experimental results show that, on average, the TF-Cache architecture reduces the dynamic power consumption up to 74.9% and 85.9% when applied to the L1 and L2 cache, respectively, for the studied applications.
Joan J. Valls, Julio Sahuquillo, Alberto Ros 0001, María Engracia Gómez
PDP4
2015 A HoL-blocking aware mechanism for selecting the upward path in fat-tree topologies
Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato
J. Supercomput.3
2015 PS-Cache: an energy-efficient cache design for chip multiprocessors
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez
J. Supercomput.4
2015 PS directory: a scalable multilevel directory cache for CMPs
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez
J. Supercomput.4
2014 FT-RUFT: A Performance and Fault-Tolerant Efficient Indirect Topology
abstract
Although performance is a key design issue of interconnection networks, fault-tolerance is becoming more important due to the large amount of components of large machines. In this paper, we focus on designing a simple indirect topology with both good performance and fault-tolerance properties. The idea is to take full advantage of the network resources consumed by the topology. To do that, starting from the RUFT topology, which is a simple UMIN topology that does not tolerate any link fault, we first duplicate injection and ejection links connecting these extra links in a particular way. The resulting topology tolerates 3 network link faults and also slightly increases performance with marginal increase in the network hardware cost. Most important, contrary to most of the available topologies, the topology is able to tolerate also faults in the links that connect to end-nodes. We also propose another topology that also duplicates network links, achieving 2x performance improvements and tolerating up to 7 network link faults. These results are better than the ones obtained by a BMIN with a similar amount of resources.
Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
PDP3
2013 PS-cache: An energy-efficient cache design for chip multiprocessors
abstract
As silicon resources become increasingly abundant, core counts grow rapidly in successive chip-multiprocessors (CMP) generations. Parallel workloads represent an important segment for current and future CMPs mainly when many-core processors are considered. Unlike multiprogrammed workloads, the accessed blocks in these workloads can be classified in two categories: private, accessed only by one core, and shared, accessed by several cores. This paper takes advantage of this classification to access only a subset of the ways on each L1 cache access, thus reducing dynamic power consumption.
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez
PACT4
2013 Temporal-Aware Mechanism to Detect Private Data in Chip Multiprocessors
abstract
Most of the data referenced by sequential and parallel applications running in current chip multiprocessors are referenced by only one thread and can be considered as private data. A lot of recent proposals leverage this observation to improve many aspects of chip multiprocessors, such as reducing coherence overhead or the access latency to distributed caches. The effectiveness of those proposals depend to a large extent on the amount of detected private data. However, the mechanisms proposed so far do not consider thread migration and the private use of data within different application phases. As a result, a considerable amount of data is not detected as private. In order to make this detection more accurate and reaching more significant improvements, we propose a mechanism that is able to account for both thread migration and private data within application phases. Simulation results for 16-core systems show that, thanks to our mechanism, the average number of pages detected as private significantly increases from 43% in previous proposals up to 74% in ours. Finally, when our detection mechanism is used to deactivate the coherence for private data in a directory protocol, our proposal improves execution time by 13% with respect to previous proposals.
Alberto Ros 0001, Blas Cuesta, María Engracia Gómez, Antonio Robles, José Duato
ICPP3
2013 Increasing the Effectiveness of Directory Caches by Avoiding the Tracking of Noncoherent Memory Blocks
abstract
A key aspect in the design of efficient multiprocessor systems is the cache coherence protocol. Although directory-based protocols constitute the most scalable approach, the limited size of the directory caches together with the growing size of systems may cause frequent evictions and, consequently, the invalidation of cached blocks, which jeopardizes system performance. Directory caches keep track of every memory block stored in processor caches in order to provide coherent access to the shared memory. However, a significant fraction of the cached memory blocks do not require coherence maintenance (even in parallel applications) because they are either accessed by just one processor or they are never modified. In this paper, we propose to deactivate the coherence protocol for those blocks that do not require coherence. This deactivation means directory caches do not have to keep track of noncoherent blocks, which reduces directory cache occupancy and increases its effectiveness. Since the detection of noncoherent blocks is carried out by the operating system, our proposal only requires minor hardware modifications. Simulation results show that, thanks to our proposal, directory caches can avoid the tracking of about 66 percent (on average) of the blocks accessed by a wide range of applications, thereby improving the efficiency of directory caches. This contributes either to shortening the runtime of parallel applications by 15 percent (on average) while keeping directory cache size or to maintaining performance while using directory caches 16 times smaller.
Blas Cuesta, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
IEEE Trans. Computers3
2012 PS-Dir: a scalable two-level directory cache
abstract
As the number of cores increases in both incoming and future chip multiprocessors, coherence protocols must address novel hardware structures in order to scale in terms of performance, power, and area. It is well known that most blocks accessed by parallel applications are private (i.e., accessed by a single core). These blocks present different directory requirements and behavior than shared blocks. Based on this fact, this paper proposes a two-level directory cache that tracks shared blocks in a small and fast first-level cache and private blocks in a larger and slower second-level cache, namely Shared and Private caches, respectively. Speed and area reasons suggest the use of eDRAM technology much dense but slower than SRAM technology for the Private cache, which in turn brings energy savings. Experimental results for a 16-core system show improvements in performance by 11.1%, in area by 25.4%, and in energy consumption by 20.5% compared to a conventional directory cache.
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez, José Duato
PACT4
2012 Towards an Efficient Fat-Tree like Topology
Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
Euro-Par3
2012 IODET: A HoL-blocking-aware Deterministic Routing Algorithm for Direct Topologies
abstract
In large parallel computers routing is a key design point to obtain the maximum possible performance out of the interconnection network. Routing can be classified into two categories depending on the number of routing options that a packet can use to go from its source to its destination. If the packet can only use a single predetermined path then the routing is deterministic, whereas if several paths are possible it is adaptive. It is a well-known fact that adaptive routing usually outperforms deterministic routing; but in this paper we take the challenge of developing a HOL-blocking-aware deterministic routing algorithm that can obtain a similar or even better performance than adaptive routing, while decreasing its implementation complexity and providing some inherent advantages to deterministic routing such as in-order delivery of packets. In this large computers regular direct topologies are widely-used, so in this paper we focus on meshes and tori.
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
ICPADS3
2012 Cache Miss Characterization in Hierarchical Large-Scale Cache-Coherent Systems
abstract
There is a growing trend towards developing large-scale cache-coherent systems by using commodity symmetric multiprocessors, which requires to extend their coherence protocol. In such systems, cache coherence transactions issued due to cache misses traverse interconnection networks with very different topologies and latencies. In this work, we perform a cache miss characterization aimed at analyzing the benefits that can be expected for a specialized coherence controller able to locally resolve cache misses, thus saving traffic across long-latency links. Results show that there is a high potential in reducing miss latency in these systems, and that this potential reduction grows as the number of nodes in the system increases. Particularly, in a system with just two boards 40% of the cache misses do not need the expensive inter-board communication. This percentage can increase up to 67.5% for an 8-board system.
Alberto Ros 0001, Blas Cuesta, María Engracia Gómez, Antonio Robles, José Duato
ISPA3
2012 A New Family of Hybrid Topologies for Large-Scale Interconnection Networks
abstract
In large supercomputers the topology of the interconnection network is a key design issue that impacts the performance and cost of the whole system. Direct topologies provide a reduced hardware cost, but as the number of dimensions is conditioned by 3D wiring restrictions, a high number of nodes per dimension is used, which increases communication latency and reduces network throughput. On the other hand, indirect topologies can provide better performance for large network sizes, but at the cost of a high amount of switches and links. In this paper we propose a new family of topologies that combines the best features of both direct and indirect topologies to efficiently connect an extremely high number of nodes. In particular, we propose an n-dimensional topology where the nodes of each dimension are connected through a small indirect topology. This combination results in a family of topologies that provides high performance, with latency and throughput figures of merit close to indirect topologies, but with a lower hardware cost. In particular, it is able to double the throughput obtained per switching element of indirect topologies. Moreover, the layout of the topology is much simpler than in indirect topologies. Indeed, its fault-tolerance degree is equal or higher than the one for direct and indirect topologies.
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
NCA3
2012 Extending Magny-Cours Cache Coherence
abstract
One cost-effective way to meet the increasing demand for larger high-performance shared-memory servers is to build clusters with off-the-shelf processors connected with low-latency point-to-point interconnections like HyperTransport. Unfortunately, HyperTransport addressing limitations prevent building systems with more than eight nodes. While the recent High-Node Count HyperTransport specification overcomes this limitation, recently launched twelve-core Magny-Cours processors have already inherited it and provide only 3 bits to encode the pointers used by the directory cache which they include to increase the scalability of their coherence protocol. In this work, we propose and develop an external device to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the directory cache. Evaluation results for systems with up to 32 nodes show that the performance offered by our solution scales with the number of nodes, enhancing the directory cache effectiveness by filtering additional messages. Particularly, we reduce execution time by 47 percent in a 32-die system with respect to the 8-die Magny-Cours configuration.
Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato
IEEE Trans. Computers4
2011 Exploiting Network-on-Chip structural redundancy for a cooperative and scalable built-in self-test architecture
abstract
This paper proposes a built-in self-test/self-diagnosis procedure at start-up of an on-chip network (NoC). Concurrent BIST operations are carried out after reset at each switch, thus resulting in scalable test application time with network size. The key principle consists of exploiting the inherent structural redundancy of the NoC architecture in a cooperative way, thus detecting faults in test pattern generators too. At-speed testing of stuck-at faults can be performed in less than 1200 cycles regardless of their size, with an hardware overhead of less than 11%.
Alessandro Strano, Crispín Gómez Requena, Daniele Ludovici, Michele Favalli, María Engracia Gómez, Davide Bertozzi
DATE5
2011 Increasing the effectiveness of directory caches by deactivating coherence for private memory blocks
abstract
To meet the demand for more powerful high-performance shared-memory servers, multiprocessor systems must incorporate efficient and scalable cache coherence protocols, such as those based on directory caches. However, the limited directory cache size of the increasingly larger systems may cause frequent evictions of directory entries and, consequently, invalidations of cached blocks, which severely degrades system performance.
Blas Cuesta, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato
ISCA3
2011 How to reduce packet dropping in a bufferless NoC
abstract
Abstract Networks on‐chip (NoCs) interconnect the components located inside a chip. In multicore chips, NoCs have a strong impact on the overall system performance. NoC bandwidth is limited by the critical path delay. Recent works show that the critical path delay is heavily affected by switch port buffer size. Therefore, by removing buffers, switch clock frequency can be increased. Recently, a new switching technique for NoCs called Blind Packet Switching (BPS) has been proposed, which is based on removing the switch port buffers. Since buffers consume a high percentage of switch power and area, BPS not only improves performance but also reduces power and area. In BPS, as there are no buffers at the switch ports, packets cannot be stopped and stored on them. If contention arises packets are dropped and later reinjected, negatively affecting performance. In order to prevent packet dropping, some techniques based on resource replication have been proposed. In this paper, we propose some alternative and complementary techniques that do not rely on resource replication. By using them, packet dropping is highly reduced. In particular, packet dropping is completely removed for a very wide network traffic range. Moreover, network throughput is increased and packet latency is reduced. Copyright © 2010 John Wiley & Sons, Ltd.
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
Concurr. Comput. Pract. Exp.2
2010 EMC2: Extending Magny-Cours coherence for large-scale servers
abstract
The demand of larger and more powerful high-performance shared-memory servers is growing over the last few years. To meet this need, AMD has recently launched the twelve-core Magny-Cours processors. They include a directory cache (Probe Filter) that increases the scalability of the coherence protocol applied by Opterons, based on coherent Hyper Transport interconnect (cHT). cHT limits up to 8 the number of nodes that can be addressed. Recent High Node Count HT specification overcomes this limitation. However, the 3-bit pointer used by the Probe Filter prevents Magny-Cours-based servers from being built beyond 8 nodes. In this paper, we propose and develop an external logic to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the Probe Filter. Evaluation results for up to a 32-node system show how the performance offered by our solution scales with the increment in the number of nodes, enhancing the Probe Filter effectiveness by filtering additional messages. Particularly, we reduce runtime by 47% in a 32-die system respect to the 8-die Magny-Cours system.
Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato
HiPC4
2010 Improved Utilization of NoC Channel Bandwidth by Switch Replication for Cost-Effective Multi-processor Systems-on-Chip
abstract
Virtual channels are an appealing flow control technique for on-chip interconnection networks (NoCs), in that they can potentially avoid deadlock and improve link utilization and network throughput. However, their use in the resource constrained multi-processor system-on-chip (MPSoC) domain is still controversial, due to their significant overhead in terms of area, power and cycle time degradation. This paper proposes a simple yet efficient approach to VC implementation, which results in more area- and power-saving solutions than conventional design techniques. While these latter replicate only buffering resources for each physical link, we replicate the entire switch and prove that our solution is counter intuitively more area/power efficient while potentially operating at higher speeds. This result builds on a well-known principle of logic synthesis for combinational circuits (the area-performance trade-off when inferring a logic function into a gate-level netlist), and proves that when a designer is aware of this, novel architecture design techniques can be conceived.
Francisco Gilabert Villamón, María Engracia Gómez, Simone Medardoni, Davide Bertozzi
NOCS2
2009 Assessing fat-tree topologies for regular network-on-chip design under nanoscale technology constraints
abstract
Most of past evaluations of fat-trees for on-chip interconnection networks rely on oversimplifying or even irrealistic architecture and traffic pattern assumptions, and very few layout analyses are available to relieve practical feasibility concerns in nanoscale technologies. This work aims at providing an in-depth assessment of physical synthesis efficiency of fat-trees and at extrapolating silicon-aware performance figures to back-annotate in the system-level performance analysis. A 2D mesh is used as a reference architecture for comparison, and a 65 nm technology is targeted by our study. Finally, in an attempt to mitigate the implementation cost of k-ary n-tree topologies, we also review an alternative unidirectional multi-stage interconnection network which is able to simplify the fat-tree architecture and to minimally impact performance.
Daniele Ludovici, Francisco Gilabert Villamón, Simone Medardoni, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, Georgi Gaydadjiev, Davide Bertozzi
DATE5
2009 FT2EI: A Dynamic Fault-Tolerant Routing Methodology for Fat Trees with Exclusion Intervals
abstract
Fault tolerance in the interconnection network of large clusters of PCs is an issue of growing importance, since their increasing size also increases the failure probability. The fat-tree topology is usually used in these machines since it has become very popular among high-speed interconnect manufacturers. This paper proposes a new distributed fault-tolerant routing methodology for fat trees. Unlike other previous proposals, it does not require additional network hardware, and its memory requirements, switch hardware, and routing delay scales up with the network size. Indeed, it nullifies only the strictly necessary paths, allowing adaptive routing through the healthy paths. The methodology is based on enhancing the interval routing scheme with exclusion intervals. Exclusion intervals are associated to each switch output port and represent the nodes that are unreachable from this port after a fault. We propose a methodology to identify the links where the exclusion intervals must be updated after a fault, the values to write on them, and a very efficient mechanism to distribute the required information through the network without stopping the system activity. Our methodology can tolerate a high number of network failures with a low degradation in performance. Moreover, it can achieve zero packet losing during the updating period.
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, J. F. D. Marin
IEEE Trans. Parallel Distributed Syst.2
2008 Reducing Packet Dropping in a Bufferless NoC
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
Euro-Par2
2008 An Efficient Switching Technique for NoCs with Reduced Buffer Requirements
abstract
Networks on chip (NoCs) communicate the components located inside a chip. Overall system performance depends on NoC performance, that is affected by several factors. One of them is the network clock frequency, imposed by the critical path delay. Recent works show that switch critical path includes buffer control logic. Consequently, by removing switch buffers, switch frequency can be doubled. In this paper, we exploit this idea, proposing a new switching technique for NoCs which requires a reduced amount of storage at the switches. It is based on replacing switch port buffers by single latches. By doing so, network cycle can be reduced, which reduces packet latency. On the other hand, power and area consumption requirements can be reduced. However, since there are no buffers at the switch ports, packets can not be stopped. Stopped packets due to contention are dropped and reinjected from their senders via negative acknowledgments. Packet dropping is strongly reduced by exploiting NoCs wiring capability.
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
ICPADS2
2008 RUFT: Simplifying the Fat-Tree Topology
abstract
The fat-tree is one of the most widely-used topologies by interconnection network manufacturers. Recently, a deterministic routing algorithm that optimally balances the network traffic in fat--trees was proposed. It can not only achieve almost the same performance than adaptive routing, but also outperforms it for some traffic patterns. Nevertheless, fat--trees require a high number of switches with a non-negligible wiring complexity. In this paper, we propose replacing the fat--tree by an unidirectional multistage interconnection network referred to as Reduced Unidirectional Fat--tree (RUFT) that uses a a simplified version of the aforementioned deterministic routing algorithm. As a consequence, switch hardware is almost reduced to the half, decreasing, in this way, power consumption, arbitration complexity, switch size, and network cost. Evaluation results show that RUFT obtains lower latency than fat--tree for low and medium traffic loads. Furthermore, in large networks, it obtains almost the same throughput than the classical fat-tree.
Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato
ICPADS3
2008 Exploring High-Dimensional Topologies for NoC Design Through an Integrated Analysis and Synthesis Framework
Francisco Gilabert Villamón, Simone Medardoni, Davide Bertozzi, Luca Benini, María Engracia Gómez, Pedro López 0001, José Duato
NOCS5
2008 Exploiting Wiring Resources on Interconnection Network: Increasing Path Diversity
abstract
On-chip networks are the answer to the growing demands for high communication performance of chip multiprocessors. These networks have a number of characteristics that make their design quite different to off-chip networks. In particular, wires are an abundant available resource inside the chip. In this paper, we explore how to organize the huge wiring capabilities available in on-chip networks. In particular, we analyze the option of distributing the wires among several parallel links connecting the same two switches. This technique is known as Space Division Multiplexing (SDM). The number of parallel sub-links and their width are two key parameters that are studied together with the relationship with the mean packet size. The paper shows that SDM is a technique to take into account in on-chip networks since it allows to highly increase the network accepted traffic at the expense of a small latency increase or even no increase. Moreover, in some networks, it allows to reduce the network hardware, providing simiar performance results, which results in a reduction in the consumption of area and power.
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
PDP2
2007 Deterministic versus Adaptive Routing in Fat-Trees
abstract
Clusters of PCs have become very popular to build high performance computers. These machines use commodity PCs linked by a high speed interconnect. Routing is one of the most important design issues of interconnection networks. Adaptive routing usually better balances network traffic, thus allowing the network to obtain a higher throughput. However, adaptive routing introduces out-of-order packet delivery, which is unacceptable for some applications. Concerning topology, most of the commercially available interconnects are based on fat-tree. Fat-trees offer a rich connectivity among nodes, making possible to obtain paths between all source-destination pairs that do not share any link. We exploit this idea to propose a deterministic routing algorithm for fat-trees, comparing it with adaptive routing in several workloads. The results show that deterministic routing can achieve a similar, and in some scenarios higher, level of performance than adaptive routing, while providing in-order packet delivery.
Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato
IPDPS3
2007 An Efficient Fault-Tolerant Routing Methodology for Fat-Tree Interconnection Networks
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato
ISPA2
2006 On the Influence of the Selection Function on the Performance of Fat-Trees
Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato
Euro-Par2
2006 FIR: An efficient routing strategy for tori and meshes
María Engracia Gómez, Pedro López 0001, José Duato
J. Parallel Distributed Comput.1
2006 A Routing Methodology for Achieving Fault Tolerance in Direct Networks
abstract
Massively parallel computing systems are being built with thousands of nodes. The interconnection network plays a key role for the performance of such systems. However, the high number of components significantly increases the probability of failure. Additionally, failures in the interconnection network may isolate a large fraction of the machine. It is therefore critical to provide an efficient fault-tolerant mechanism to keep the system running, even in the presence of faults. This paper presents a new fault-tolerant routing methodology that does not degrade performance in the absence of faults and tolerates a reasonably large number of faults without disabling any healthy node. In order to avoid faults, for some source-destination pairs, packets are first sent to an intermediate node and then from this node to the destination node. Fully adaptive routing is used along both subpaths. The methodology assumes a static fault model and the use of a checkpoint/restart mechanism. However, there are scenarios where the faults cannot be avoided solely by using an intermediate node. Thus, we also provide some extensions to the methodology. Specifically, we propose disabling adaptive routing and/or using misrouting on a per-packet basis. We also propose the use of more than one intermediate node for some paths. The proposed fault-tolerant routing methodology is extensively evaluated in terms of fault tolerance, complexity, and performance.
María Engracia Gómez, Nils Agne Nordbotten, José Flich, Pedro López 0001, Antonio Robles, José Duato, Tor Skeie, Olav Lysne
IEEE Trans. Computers1
2005 A Memory-Effective Fault-Tolerant Routing Strategy for Direct Interconnection Networks
abstract
High-performance interconnection networks are crucial in massively parallel computers. Routing is one of the most important design issues of interconnection networks. Moreover, the huge amount of hardware of these machines makes fault-tolerance another important design issue. In this paper, we propose a mechanism that combines scalable routing and fault-tolerance for commercial switches to build direct regular topologies, which are the topologies used in large machines. The hardware required is not complex. Furthermore, it allows a high degree of fault-tolerance inflicting a minimal decrease of performance
María Engracia Gómez, Pedro López 0001, José Duato
ISPDC1
2004 A New Adaptive Fault-Tolerant Routing Methodology for Direct Networks
María Engracia Gómez, José Duato, José Flich, Pedro López 0001, Antonio Robles, Nils Agne Nordbotten, Tor Skeie, Olav Lysne
HiPC1
2004 An Effective Fault-Tolerant Routing Methodology for Direct Networks
abstract
Current massively parallel computing systems are being built with thousands of nodes, which significantly affect the probability of failure. M. E. Gomex proposed a methodology to design fault-tolerant routing algorithms for direct interconnection networks. The methodology uses a simple mechanism: for some source-destination pairs, packets are first forwarded to an intermediate node, and later, from this node to the destination node. Minimal adaptive routing is used along both subpaths. For those cases where the methodology cannot find a suitable intermediate node, it combines the use of intermediate nodes with two additional mechanisms: disabling adaptive routing and using misrouting on a per-packet basis. While the combination of these three mechanisms tolerates a large number of faults, each one requires adding some hardware support in the network and also introduces some overhead. In this paper, we perform an in-depth detailed analysis of the impact of these mechanisms on network behaviour. We analyze the impact of the three mechanisms separately and combined. The ultimate goal of this paper is to obtain a suitable combination of mechanisms that is able to meet the trade-off between fault-tolerance degree, routing complexity, and performance.
María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, José Duato, Nils Agne Nordbotten, Olav Lysne, Tor Skeie
ICPP1
2004 A Fully Adaptive Fault-Tolerant Routing Methodology Based on Intermediate Nodes
Nils Agne Nordbotten, María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, Tor Skeie, Olav Lysne, José Duato
NPC2
2002 Evaluation of Routing Algorithms for InfiniBand Networks (Research Note)
María Engracia Gómez, José Flich, Antonio Robles, Pedro López 0001, José Duato
Euro-Par1
2000 A new approach in the analysis and modeling of disk access patterns
abstract
While in previous work we have demonstrated that disk arrival patterns are consistent with self-similarity and have provided a physical explanation for the self-similar phenomenon in disk arrival patterns, the authors now deal with the analysis and modeling of disk access patterns. We provide visual and mathematical evidence showing that the same bursty behavior observed in the time series can also be observed in the spatial series formed by the starting address of the accessed blocks. Moreover, we demonstrate that the clustering of accesses in areas of the disk has self-similar characteristics. Next, we demonstrate that the observed behavior can be explained using the same self-similar property used in the analysis of arrival patterns and applying the same physical explanation.
María Engracia Gómez, Vicente Santonja
ISPASS1
2000 A New Approach in the Modeling and Generation of Synthetic Disk Workload
abstract
Shows a new approach to generate synthetic disk workload. The work presented is based on previous results obtained from the analysis of real disk traces. The proposed disk workload generation model can capture the heavy-tailed behavior of real disk workload, a critical feature to reproduce disk subsystem congestion. The generator provides synthetic workload much more accurately than commonly-used models. Since the workload plays a critical role in performance evaluations, having a more accurate disk workload generator is important for storage researchers in order to obtain fair and unbiased performance predictions.
María Engracia Gómez, Vicente Santonja
MASCOTS1
1999 Analysis of Self-Similarity in I/O Workload Using Structural Modeling
abstract
Demonstrates that disk-level I/O requests are self-similar in nature. We show evidence (both visual and mathematical) that I/O accesses are consistent with self-similarity. For this analysis, we have used two sets of disk activity traces collected from various systems over different periods of time. In addition to studying the aggregated I/O workload that is directed to the storage system, we perform a structural modeling of the workload in order to understand the underlying causes that produce the observed self-similarity. This structural modeling shows that self-similar behavior can be explained by combining two different approaches: the on/off source model and Cox's model. The former applies to those processes that remain active during the whole trace, while the latter applies to sources that show a very short activity time.
María Engracia Gómez, Vicente Santonja
MASCOTS1