EDBT 2026 Demo / reviewers in the wild / expert
Julio Sahuquillo
dblp:59/6649 · also Julio Sahuquillo Borrás
· DBLP profile ↗
127ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0001-8630-4846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 91 · 2 first-author · 18 since 2021Computer networks · 8Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SYNPA: understanding the impact of the interference modeling on thread-to-core allocation policies for SMT ARM processorsabstractModern high-performance servers increasingly rely on Simultaneous Multithreading (SMT) processors to enhance throughput with minimal area overhead. However, SMT architectures introduce inter-application interference, often resulting in degraded performance for individual applications. To address this issue, interference-aware thread-to-core (T2C) allocation policies are essential. This paper explores the design and implementation of such policies using real performance counters on ARM processors. We introduce the Instructions and Stalls Cycles (ISC) stack—a simple yet effective model for characterizing application behavior and identifying synergistic thread pairings. Building on our previous work, SYNPA, we improve the accuracy of the model by accounting for horizontal waste (that is, unused dispatch slots) and proposing methods to address limitations in ARM’s Performance Monitoring Unit (PMU), which prevent complete attribution of processor cycles. These enhancements result in a family of SYNPA schedulers, each based on a different ISC stack variant. Detailed discussions are provided on the pros and cons that researchers typically face when building a performance stack on commercial processors. These analyses are intended to assist researchers in their work. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
J. Supercomput. | 5 |
| 2026 | WAPA: A Microarchitecture- and Workload-Agnostic Universal SMT Scheduler
Marta Navarro 0001, Vicent Pallardó-Julià, Lucia Pons, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | WAPA: A Workload-Agnostic CPI-Based Thread-to-Core Allocation Policy
Marta Navarro 0001, Vicent Pallardó-Julià, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
Euro-Par (1) | 5 |
| 2025 | Power, energy, and performance analysis of single- and multi-threaded applications in the ARM ThunderX2abstractEnergy efficiency has been a major concern in data centers, and the problem is exacerbated as its size continues to rise. However, the lack of tools to measure and handle this energy at a fine granularity (e.g., processor core or last-level cache) has translated into slow research advances in this topic. Understanding where (i.e., which components) and when (the point in time) energy consumption translates into minor performance improvements is of paramount importance to design any energy-aware scheduler. This paper characterizes the relationship between energy consumption and performance in a 28-core ARM ThunderX2 processor for both single-threaded and multi-threaded applications. This paper shows that single-threaded applications with high CPU activity maintain their performance in spite of the inter-application interference at shared resources, but this comes at the expense of higher power consumption. Conversely, applications that heavily utilize the L3 cache and memory consume less power but suffer significant performance degradation as interference levels rise. In contrast, multi-threaded applications show two distinct behaviors. On the one hand, some of them experience significant performance gains when they execute in a higher number of cores with more threads, which outweighs the increase in power consumption, leading to high energy efficiency. Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
J. Parallel Distributed Comput. | 4 |
| 2025 | Advanced resource management: A hands-on master course in HPC and cloud computingabstractResource management has become a major concern in dealing with performance and fairness in recent computing servers, including a wide variety of shared resources. To achieve high-performing and efficient systems, both hardware and software engineers must be thoroughly trained in effective resource management techniques. This paper introduces the GRE master course (Spanish acronym for Resource Management and Performance Evaluation in Cloud and High-Performance Workloads), which is being offered since Fall 2023. The course is taught by instructors with broad research expertise in resource management and performance evaluation. Subjects covered in this course include workload characterization, state-of-the-art resource management approaches, and performance evaluation tools and methodologies used in production systems. Management techniques are studied both in the context of HPC and cloud computing, where resource efficiency is becoming a primary concern. To enhance the learning experience, the course integrates theoretical concepts with a wide set of hands-on tasks carried out on recent real platforms. A real cloud virtualized environment is mimicked using typical software deployed in production systems such as Proxmox Virtual Environment. Students learn to use tools such as Linux Perf and Intel Vtune Profiler, which are commonly employed by researchers and practitioners to carry out typical tasks like performance bottleneck analysis from a microarchitectural perspective. Overall, the GRE course provides students with a solid foundation and skills in resource management by addressing current hot topics both in the industry and academia. Student satisfaction and learning outcomes prove the success of the GRE course and encourage us to continue in this direction. Lucia Pons, Salvador Petit, Julio Sahuquillo |
J. Parallel Distributed Comput. | 3 |
| 2025 | Dual Fast-Track Cache: Organizing Ring-Shaped Racetracks to Work as L1 CachesabstractStatic Random-Access Memory (SRAM) is the fastest memory technology and has been the common design choice for implementing first-level (L1) caches in the processor pipeline, where speed is a key design issue that must be fulfilled. On the contrary, this technology offers much lower density compared to other technologies like Dynamic RAM, limiting L1 cache sizes of modern processors to a few tens of KB.This paper explores the use of slower but denser Domain Wall Memory (DWM) technology for L1 caches. This technology provides slow access times since it arranges multiple bits sequentially in a magnetic racetrack. To access these bits, they need to be shifted in order to place them under a header. A 1-bit shift usually takes one processor cycle, which can significantly hurt the application performance, making this working behavior inappropriate for L1 caches.Based on the locality (temporal and spatial) principles exploited by caches, this work proposes the Dual Fast-Track Cache (Dual FTC) design, a new approach to organizing a set of racetracks to build set-associative caches. Compared to a conventional SRAM cache, Dual FTC enhances storage capacity by a factor of 5 while incurring minimal shifting overhead, thereby rendering it a practical and appealing solution for L1 cache implementations.Experimental results show that the devised cache organization is as fast as an SRAM cache for 78% and 86% of the L1 data cache hits and L1 instruction cache hits, respectively (i.e., no shift is required). Consequently, due to the larger L1 cache capacities, significant system performance gains (by 22% on average) are obtained under the same silicon area. Alejandro Valero, Vicente Lorente, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 4 |
| 2024 | Optimization of One-to-Many Communication Primitives for Dragonfly TopologiesabstractCollective communication primitives (CCPs), such as multicast and broadcast, are essential for many parallel and distributed applications. In response, this study compares a topology-oblivious algorithm underlying the implementation of CCPs in standard instances of MPI with two topology-aware implementations, based on the LLF and GLF algorithms, and an ideal hardware-assisted approach. By using real scientific applications, instead of synthetic traffic, our study reveals workload-dependent performance variations among CCP implementations; and highlights the importance of CCP algorithm selection in optimizing application performance in supercomputing environments. Jose Duro, Adrián Castelló 0001, María Engracia Gómez, Julio Sahuquillo, Enrique S. Quintana-Ortí |
ICPADS | 4 |
| 2024 | SYNPA: SMT Performance Analysis and Allocation of Threads to Cores in ARM ProcessorsabstractSimultaneous multithreading processors improve throughput over single-threaded processors thanks to sharing internal core resources among instructions from distinct threads. However, resource sharing introduces inter-thread interference within the core, which has a negative impact on individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the interference effects depends on the competing co-runners sharing the core. Thus, it can be mitigated by applying a thread-to-core allocation policy that smartly selects applications to be run in the same core to minimize their interference.This paper presents SYNPA, a simple approach that dynamically allocates threads to cores in an SMT processor based on their run-time dynamic behavior. The approach uses a regression model to select synergistic pairs to mitigate intra-core interference. The main novelty of SYNPA is that it uses just three variables collected from the performance counters available in current ARM processors at the dispatch stage. Experimental results show that SYNPA outperforms the default Linux scheduler by around 36%, on average, in terms of turnaround time in 8-application workloads combining frontend-bound and backend-bound benchmarks. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IPDPS | 5 |
| 2024 | Characterizing Power and Performance Interference Scalability in the 28-core ARM ThunderX2abstractNowadays, energy efficiency is a major concern in any type of processor-based device, ranging from processor servers to supercomputers, including mobile battery-fed devices. In recent years, systems based on the ARM architecture, tra-ditionally better suited for mobile and embedded systems, have increased their market share in segments commonly occupied by x86 processors. This growth is due, to some extent, to the excellent energy efficiency shown by high-performance ARM processors. Designing software and hardware energy-efficient systems requires a sound knowledge of the relationship among three main axes: component activity, power, and inter-application interference at the shared resources. This problem aggravates in many-core processors, which are ubiquitous in high-performance servers. This paper characterizes the aforementioned axes in a 28-core ARM Thunder X2 processor. Experimental results show that the performance of highly-scalable single-threaded applications is sustained regardless of the number of applications and interference introduced at the shared resources at the cost of increasing power. In contrast, low-scalable applications require much less power and experience a huge performance degradation as their number increases. Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PDP | 4 |
| 2024 | A modular approach to build a hardware testbed for cloud resource management research
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo |
J. Supercomput. | 5 |
| 2023 | Thread-to-Core Allocation in ARM Processors Building Synergistic PairsabstractSimultaneous multithreading (SMT) processors can present significant throughput improvements over single-threaded (ST) processors thanks to sharing internal core resources among instructions executing from multiple threads. However, resource sharing introduces inter-thread interference within the core, which negatively impacts individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the intra-core interference on performance depends on the applications co-running in the same core. A thread scheduler can help reduce this effect by smartly selecting the pairs of applications that should run on each SMT core. This paper presents SYNPA, a simple approach that dynamically allocates threads to SMT cores based on their run-time dynamic behavior. SYNPA uses a regression model to select synergistic pairs to mitigate intra-core interference. Results show that SYNPA outperforms the default Linux scheduler by around 35%, on average, in terms of turnaround time when running 8-application workloads combining frontend-bound and backend-bound applications. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PACT | 5 |
| 2023 | Dynamic Allocation of Processor Cores to Graph Applications on Commodity ServersabstractGraph processing is increasingly adopted to solve problems that span many application domains, including scientific computing, social networks, and big-data analytics. These applications present particular features (huge working sets and irregular scalability) that make the default Linux scheduler, which adopts a time-sharing policy to provide a fair scheduler, perform poorly when co-locating multiple graph applications in the same processor. This work focuses on maximizing processor utilization, which is a major concern of current data centers. To this end, we propose AFAIR, a flexible scheduling policy that allocates multiple graph applications on the same processor and assigns a fraction of the cores exclusively to each application instead of sharing them. Moreover, AFAIR dynamically adds/removes cores to the running applications, adapting the number of threads used for parallel execution to balance memory load. This allows AFAIR to achieve almost perfect fairness, on average 95%. Lucia Pons, Julio Sahuquillo, Timothy M. Jones 0001 |
PACT | 2 |
| 2023 | Stratus: A Hardware/Software Infrastructure for Controlled Cloud ResearchabstractCloud systems deploy a wide variety of shared resources and host a large number of tenant applications. To perform cloud research, a small experimental platform is commonly used, which hides the huge system complexity and provides flexibility. Despite being simpler, this platform should include the main cloud system components (hardware and software) to provide representative results. A wide set of platforms have spread in recent years; however, most of them only include a major cloud component or lack the deployment of virtual machines (VMs) to provide isolation. This paper presents Stratus, an experimental platform that is currently being used to carry out cloud research. To the best of our knowledge, Stratus is the only platform that jointly provides three main features: uses VMs to isolate tenant applications, deploys the three types of cloud nodes (server, client, and storage), and manages all main shared system resources (CPUs, LLC space, memory, network, and disk bandwidth). Moreover, Stratus implements a software manager to ease the research and aid the design of QoS-aware policies. The manager integrates three main functionalities: management and control of the execution of VMs and running applications, monitoring of hardware performance counters and system resource utilization, and partitioning of the main shared system resources by using technologies available in commercial processors. Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Chaoyi Huang, Julio Sahuquillo |
PDP | 6 |
| 2023 | Cloud White: Detecting and Estimating QoS Degradation of Latency-Critical Workloads in the Public CloudabstractThe increasing popularity of cloud computing has forced cloud providers to build economies of scale to meet the growing demand. Nowadays, data-centers include thousands of physical machines, each hosting many virtual machines (VMs), which share the main system resources, causing interference that can significantly impact on performance. Frequently, these data-centers run latency-critical workloads, whose performance is determined by tail latency, which is very sensitive to the interference of co-running workloads. To prevent QoS violations, cloud providers adopt overprovisioning strategies but they reduce the server utilization and increase the costs. A mechanism that accurately estimates performance degradation dynamically in a production system would allow cloud providers to improve the servers’ utilization. In this work we propose Cloud White, an approach that is able to detect the inter-VM interference in scenarios with multiple co-located latency-critical VMs and estimate the performance degradation using multi-variable regression models. Unlike previous proposals, Cloud White is built taking into account the limitations of a public cloud production system. Experimental results show that Cloud White is able to estimate performance degradation with a small overall prediction error of 5%. Lucia Pons, Josué Feliu, Julio Sahuquillo, María Engracia Gómez, Salvador Petit, Julio Pons, Chaoyi Huang |
Future Gener. Comput. Syst. | 3 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 67 |
| 2022 | Cache-Poll: Containing Pollution in Non-Inclusive Caches Through Cache PartitioningabstractCurrent server processors have redistributed the cache hierarchy space over previous generations. The private L2 cache has been made larger and the shared last level caches (LLC) smaller but designed as non-inclusive to reduce the number of replicated blocks. As a result, the new organization shrinks the per-core cache area. Lucia Pons, Julio Sahuquillo, Salvador Petit, Julio Pons |
ICPP | 2 |
| 2022 | Fast-track cache: a huge racetrack memory L1 data cacheabstractFirst-level (L1) caches have been traditionally implemented with Static Random-Access Memory (SRAM) technology, since it is the fastest memory technology, and L1 caches call for tight timing constraints in the processor pipeline. However, one of the main downsides of SRAM is its low density, which prevents L1 caches to improve their storage capacity beyond a few tens of KB. On the other hand, the recent Domain Wall Memory (DWM) technology overcomes such a constraint by arranging multiple bits in a magnetic racetrack, and sharing a header to access those bits. Accessing a bit requires a shift operation to align the target bit under the header. Such shifts increase the final access latency, which is the main reason why DWM has been mostly used to implement slow last-level caches. Hugo Tárrega, Alejandro Valero, Vicente Lorente, Salvador Petit, Julio Sahuquillo |
ICS | 5 |
| 2022 | A Neural Network to Estimate Isolated Performance from Multi-Program ExecutionabstractWhen multiple applications are running on a platform with shared resources like multicore CPUs, the behaviour of the running application can be altered by the co-runners. In this case, the system resources need to be managed (e.g. by repartitioning the cache space, re-schedule applications in distinct cores, modifying the prefetcher configuration, etc.) to reduce the inter-application interference in order to minimize the performance losses over isolated execution. In this context, a main challenge in different computing scenarios like the public cloud or soft real-time systems is knowing the performance impact of a given management action on each application with respect to its isolated execution. With this aim, in this work we present a neural network-based approach that estimates the performance an application would have had in isolation from multi-program executions. Experimental results show that the proposal dynamically adapts to changes in application behavior. On average, the predicted performance presents an error deviation by 11.7% and 2.3% for MAPE and MSE respectively. Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PDP | 5 |
| 2022 | Effect of Hyper-Threading in Latency-Critical Multithreaded Cloud Applications and Utilization Analysis of the Major System ResourcesabstractMultithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main system resources is a major concern for cloud providers in order to devise virtualization strategies that improve the system efficiency. With this aim, this paper first characterizes the impact of QPS on tail latency, analyzing different scenarios varying the number of threads and the thread-to-core allocation (single-task and multi-task execution) policy. The characterization study reveals that the performance of some applications does not scale with the number of threads, and the performance of some others is insensitive to the Hyper-Threading technology, so they can be allocated in less physical cores and improve system utilization. Identifying these applications, however, at run-time is challenging. Despite identifying these applications at run-time is challenging, this paper shows that they can be successfully detected at run-time by analyzing the utilization trend of the major system resources. In addition to CPU, we have also studied how assigning the share of each application of other major shared system resources impacts on performance. We outline considerations cloud providers should take into account to improve performance and resource utilization. Lucia Pons, Josué Feliu, José Puche, Chaoyi Huang, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo |
Future Gener. Comput. Syst. | 8 |
| 2022 | VMT: Virtualized Multi-Threading for Accelerating Graph Workloads on Commodity ProcessorsabstractModern-day graph workloads operate on huge graphs through pointer chasing which leads to high last-level cache (LLC) miss rates and limited memory-level parallelism (MLP). Simultaneous Multi-Threading (SMT) effectively hides the memory access latencies for multi-threaded graph workloads provided that sufficient threads are supported in hardware. Unfortunately, providing a sufficiently large number of physical threads incurs an unjustifiably high hardware cost for commodity SMT processors which typically implement only two physical hardware threads. Ideally, we would like to achieve aggressive-SMT performance when running graph workloads on modest commodity processors. In this paper, we propose Virtualized Multi-Threading (VMT), a low-overhead multi-threading paradigm for accelerating graph workloads on commodity processors. Unlike prior multi-threading paradigms, VMT virtualizes both the physical hardware threads and the architecture state: VMT maps a large number of logical software threads to a small number of physical hardware threads, while maintaining the architecture state of the logical threads in the processor's cache hierarchy. Implemented on top of a quad-core 2-way SMT processor, VMT achieves an average speedup of 1.74× for a set of representative graph workloads, while incurring minimal hardware cost (195 bytes per core to support up to 32 logical threads). VMT's low hardware cost paves the way for implementation in commodity processors. Josué Feliu, Ajeya Naithani, Julio Sahuquillo, Salvador Petit, Moinuddin K. Qureshi, Lieven Eeckhout |
IEEE Trans. Computers | 3 |
| 2022 | DeepP: Deep Learning Multi-Program Prefetch Configuration for the IBM POWER 8abstractCurrent multi-core processors implement sophisticated hardware prefetchers, that can be configured by application (PID), to improve the system performance. When running multiple applications, each application can present different prefetch requirements, hence different configurations can be used. Setting the optimal prefetch configuration for each application is a complex task since it does not only depend on the application characteristics but also on the interference at the shared memory resources (e.g., memory bandwidth). In his paper, we proposeDeepP, a deep learning approach for the IBM POWER8 that identifies at run-time the best prefetch configuration for each application in a workload. To this end, the neural network predicts the performance of each application under the studied prefetch configurations by using a set of performance events. The prediction accuracy of the network is improved thanks to a dynamic training methodology that allows learning the impact of dynamic changes of the prefetch configuration on performance. At run-time, the devised network infers the best prefetch configuration for each application and adjusts it dynamically. Experimental results show that the proposed approach improves performance, on average, by 5.8%, 6.7%, and 15.8% compared to the default prefetch configuration across different 6-, 8-, and 10-application workloads, respectively. Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Computers | 5 |
| 2020 | Impact of the Array Shape and Memory Bandwidth on the Execution Time of CNN Systolic ArraysabstractThe use of Convolutional Neural Networks (CNN) has experienced a huge rise over the last recent years and its popularity has increased exponentially, mainly due to its application both for image recognition and certain applications related to artificial intelligence. The new applications of CNN request computing demands that are difficult to address by conventional processors.As a consequence, accelerators -both prototypes and commercial products- focusing on CNN computation have been proposed. Among these accelerators, those based on systolic arrays have acquired a special relevance; some examples are the Google's TPU and Eyeriss.Current research has focused on regular squared systolic arrays and most existing work assumes that there is enough memory bandwidth to feed the systolic array with input data. In this paper we explore the design of non-squared systolic arrays and address the impact of the memory bandwidth from a performance perspective.This work makes two main contributions. First, we found that some workloads with non-squared arrays achieve similar performance to systolic arrays twice as large, which can translate in area and/or energy benefits.Second, we present a performance comparison varying the main memory bandwidth for current DRAM devices. The analysis reveals that main memory bandwidth has a great impact on performance and that the decision of which technology use is key for the system performance. For the 64x64 array size it is necessary to use HBM2 memory to avoid the slowdown that would introduce cheaper technologies (e.g. DDR5 and DDR4). Eduardo Yago, Pau Castelló, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
DSD | 5 |
| 2020 | An efficient cache flat storage organization for multithreaded workloads for low power processors
José Puche, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
Future Gener. Comput. Syst. | 4 |
| 2020 | Thread Isolation to Improve Symbiotic Scheduling on SMT Multicore ProcessorsabstractResource sharing is a critical issue in simultaneous multithreading (SMT) processors as threads running simultaneously on an SMT core compete for shared resources. Symbiotic job scheduling, which co-schedules applications with complementary resource demands, is an effective solution to maximize hardware utilization and improve overall system performance. However, symbiotic job scheduling typically distributes threads evenly among cores, i.e., all cores get assigned the same number of threads, which we find to lead to sub-optimal performance. In this paper, we show that asymmetric schedules (i.e., schedules that assign a different number of threads to each SMT core) can significantly improve performance compared to symmetric schedules. To leverage this finding, we propose thread isolation, a technique that turns symmetric schedules into asymmetric ones yielding higher overall system performance. Thread isolation identifies SMT-adverse applications and schedules them in isolation on a dedicated core to mitigate their sharp performance degradation under SMT. Our experimental results on an IBM POWER8 processor show that thread isolation improves system throughput by up to 5.5 percent compared to a state-of-the-art symmetric symbiotic job scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, Lieven Eeckhout |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Bandwidth-Aware Dynamic Prefetch Configuration for IBM POWER8abstractAdvanced hardware prefetch engines are being integrated in current high-performance processors. Prefetching can boost the performance of most applications, however, the induced bandwidth consumption can lead the system to a high contention for main memory bandwidth, which is a scarce resource in current multicores. In such a case, the system performance can be severely damaged. This article characterizes the applications’ behavior in an IBM POWER8 machine, which presents many prefetch settings, varying the bandwidth contention. The study reveals that the best prefetch setting for each application depends on the main memory bandwidth availability, that is, it depends on the co-running applications. Based on this study, we propose Bandwidth-Aware Prefetch Configuration (BAPC) a scalable adaptive prefetching algorithm that improves the performance of multi-program workloads. BAPC increases the performance of the applications in a 12, 15, and 16 percent of 6-, 8-, and 10-application workloads over the IBM POWER8 default configuration. In addition, BAPC reduces bandwidth consumption in 39, 42, and 45 percent, respectively. Carlos Navarro, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Phase-Aware Cache Partitioning to Target Both Turnaround Time and System PerformanceabstractThe Last Level Cache (LLC) plays a key role in the system performance of current multi-cores by reducing the number of long latency main memory accesses. The inter-application interference at this shared resource, however, can lead the system to undesired situations regarding performance and fairness. Recent approaches have successfully addressed fairness and turnaround time (TT) in commercial processors. Nevertheless, these approaches must face sustaining system performance, which is challenging. This work makes two main contributions. LLC behaviors regarding cache performance, data reuse and cache occupancy, that adversely impact on the final performance are identified. Second, based on these behaviors, we propose the Critical-Phase Aware Partitioning Approach (CPA), which reduces TT while sustaining (and even improving) IPC by making an effective use of the LLC space. Experimental results show that CPA outperforms CA, Dunn and KPart state-of-the-art approaches, and improves TT (over 40 percent in some workloads) over Linux default behavior while sustaining or even improving IPC by more than 3 percent in several mixes. Lucia Pons, Julio Sahuquillo, Vicent Selfa, Salvador Petit, Julio Pons |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Foreword to the Special Issue on Processors, Interconnects, Storage, and Caches for Exascale SystemsabstractExascale computing constitutes nowadays a significant challenge both for the academia and the industry. Although traditional computer systems continue to make important advances, achieving exascale computing requires mass customization. With this aim, several ongoing research projects are focusing on different architectural (computing boards or nodes, interconnects, storage, etc) issues of future exascale systems. Most of them devise heterogeneous computing boards consisting of CPUs (high performance and/or low power), FPGAs, GPUs, etc, sharing a common memory hierarchy. In this context, efficient intra- and inter-board within the same rack and inter-rack interconnect with the memory hierarchies are required. Also, performance and reliability design constraints for exascale storage systems rise significant challenges for HPC system designers. High performance I/O must be also faced because storing and retrieving such large amounts of data can greatly affect the overall performance of applications. Finally, it is important to characterize the demands that exascale applications exert on the different components of an exascale system. The goal of this special issue is to promote research on all of these aspects related to exascale computing. Six papers that address several components of exascale systems were carefully selected from open submissions. The six manuscripts included in this special issue cover different aspects of an exascale system. Lant et al1 present a network interface architecture and networking infrastructure, designed to sit inside the FPGA fabric of a cutting-edge heterogeneous MPSoC (Multi-Processor System-on-Chip), enabling networks of these devices to communicate within both a distributed and shared memory context, with reduced need for costly software networking system calls. This work presents in detail the factors that influenced the implementation and system prototype-based upon the use of Xilinx Zynq Ultrascale+ and discusses the main design decisions and implementation challenges. Crespo et al2 emphasize the need of interconnect technologies alternative to the classical electrical one. This work focuses on silicon photonics and highlights practical challenges that must be met to enable the adoption of this technology in building efficient, extreme-scale interconnection networks. In particular, they show that signal loss sources, suffered mainly due to waveguide crossings and propagation, play a critical role in photonic exascale network designs as they constrain the ability to perform data transmission in an effective as network size increases. Also focused on the interconnection network of an exascale system, Duro et al3 conduct an extensive simulation study using realistic photonic network configurations with synthetic and realistic traffic and show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. The study is performed from an architectural perspective, and the authors state that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Piernas and González-Férez4 address the scalability of file systems aimed to exascale systems. In particular, they describe how they have implemented the support for data objects in their previously proposed Fusion Parallel File System (FPFS). They show that the utilization of a unified data and metadata server (an enhanced object-based storage device or OSD+) provides FPFS with a competitive advantage over other file systems like Lustre or OrangeF, which brings higher performance in some file operations. Metadata-intensive workloads are used to stress the network traffic and analyze the scalability of the file systems. Pascual et al5 investigate alternatives for the storage subsystem of a novel exascale-capable system with special emphasis on how allocation strategies would affect the overall performance. They consider several aspects of data-aware allocation (such as the effect of spatial and temporal locality, the affinity of data to storage sources, and the network-level traffic prioritization for different types of flows) and show that scheduling policies exposing data-locality information can be essential for the appropriate utilization of future large-scale systems. They also found that the distributed storage system they implement can outperform traditional SAN architectures, even with a much smaller (in terms of I/O servers) back-end. Finally, Castro et al6 present an energy study on critical parameters for the deployment of CNNs on flagship image and video applications, ie, object recognition and people identification by gait, respectively. Their experimental results on a multi-GPU server endowed with twin Maxwell and twin Pascal Titan X GPUs demonstrate that energy correlates with performance and that Pascal may have up to 40% gains versus Maxwell. Larger batch sizes extend performance gains and energy savings but accuracy must be watched, which sometimes shows a preference for small batches. The manuscripts presented in this special issue provide insights into several cutting-edge aspects of exascale computing. We believe that the main contributions presented in these manuscripts are timely and important. We hope that readers can benefit from these research manuscripts and contribute to these rapidly growing areas. Manuel E. Acacio is a Full Professor of computer architecture and technology at the University of Murcia, Spain. He obtained his PhD degree in Computer Science in March 2003. Before, in the summer of 2002, he worked as a summer intern at IBM TJ Watson, Yorktown Heights (NY). Currently, Prof. Acacio leads the Computer Architecture & Parallel Systems (CAPS) research group at the University of Murcia. He is author of more than 100 papers in refereed international conferences and journals. As well, he has served as a committee member of important conferences, ICPP and IPDPS among others. His research interests are focused on the architecture of multiprocessor systems. From April 2011 to April 2015, Prof. Acacio served as an associate editor of IEEE Transactions on Parallel and Distributed Systems International Journal; since August 2016, he is a member of the editorial board of MPDI Computers Int'l Journal; and more recently, since September 2018, he serves as an academic editor in the editorial board of Hindawi Scientific Programming journal. He is also a member of the board of distinguished reviewers of ACM Transactions on Architecture and Code Optimization Int'l Journal since May 2014. Julio Sahuquillo is a Full Professor with the Department of Computer Engineering at the Universitat Politècnica de València. He has enjoyed a postdoctoral research stay with Prof Antonio Gónzalez, former director of Intel Barcelona. He has taught several courses on computer organization and architecture. He has co-authored more than 150 refereed conference and journal papers. His current research interests include multi- and manycore processors, memory hierarchy design, cache coherence, GPU architecture, resource management in the cloud, and architecture-aware scheduling. In these topics, he has advised more than 10 PhD Theses and has been the principal investigator of competitive Spanish domestic projects, international European projects, and projects with international companies. He has participated in the organization of about 30 conferences (HPCC, Euro-Par, HPCS, etc.) in different positions: Publicity Chair, Local Organizing Committee, Workshop Co-Chair, and Special Session Co-Chair. He participates assiduously in the PC of major Computer Architecture conferences. He is a member of the IEEE and the IEEE Computer Society. The guess editors would like to thank all the authors who made valuable contributions to this special issue. We also thank the reviewers for their detailed review reports that have helped to further enhance the manuscripts originally submitted. Finally, we would like to express our sincere gratitude to Prof Geoffrey Fox, the editor-in-chief, for having provided us with the opportunity to edit this special issue in the international journal of Concurrency and Computation: Practice and Experience, as well as for his assistance throughout all the review process. Manuel E. Acacio, Julio Sahuquillo |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Modeling and analysis of the performance of exascale photonic networksabstractSummary Photonics technology has become a promising and viable alternative for both on‐chip and off‐chip interconnection networks of future Exascale systems. Nevertheless, this technology is not mature enough yet in this context, so research efforts focusing on photonic networks are still required to achieve realistic suitable network implementations. In this regard, system‐level photonic network simulators can help guide designers to assess the multiple design choices. Most current research is done on electrical network simulators, whose components work widely different from photonics components. In this work, we summarize and compare the working behavior of both technologies which includes the use of optical routers, wavelength‐division multiplexing and circuit switching among others. After implementing them into a well‐known simulation framework, an extensive simulation study has been carried out using realistic photonic network configurations with synthetic and realistic traffic. Experimental results show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. Our study also reveals that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Jose Duro, Jose Antonio Pascual, Salvador Petit, Julio Sahuquillo, María Engracia Gómez |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Efficient Management of Cache Accesses to Boost GPGPU Memory Subsystem PerformanceabstractTo support the massive amount of memory accesses that GPGPU applications generate, GPU memory hierarchies are becoming more and more complex, and the Last Level Cache (LLC) size considerably increases each GPU generation. This paper shows that counter-intuitively, enlarging the LLC brings marginal performance gains in most applications. In other words, increasing the LLC size does not scale neither in performance nor energy consumption. We examine how LLC misses are managed in typical GPUs, and we find that in most cases the way LLC misses are managed are precisely the main performance limiter. This paper proposes a novel approach that addresses this shortcoming by leveraging a tiny additional Fetch and Replacement Cache-like structure (FRC) that stores control and coherence information of the incoming blocks until they are fetched from main memory. Then, the fetched blocks are swapped with the victim blocks (i.e., selected to be replaced) in the LLC, and the eviction of such victim blocks is performed from the FRC. This approach improves performance due to three main reasons: i) the lifetime of blocks being replaced is enlarged, ii) the main memory path is unclogged on long bursts of LLC misses, and iii) the average LLC miss latency is reduced. The proposal improves the LLC hit ratio, memory-level parallelism, and reduces the miss latency compared to much larger conventional caches. Moreover, this is achieved with reduced energy consumption and with much less area requirements. Experimental results show that the proposed FRC cache scales in performance with the number of GPU compute units and the LLC size, since, depending on the FRC size, performance improves ranging from 30 to 67 percent for a modern baseline GPU card, and from 32 to 118 percent for a larger GPU. In addition, energy consumption is reduced on average from 49 to 57 percent for the larger GPU. These benefits come with a small area increase (by 7.3 percent) over the LLC baseline. Francisco Candel, Alejandro Valero, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 4 |
| 2019 | An Aging-Aware GPU Register File Design Based on Data RedundancyabstractNowadays, GPUs sit at the forefront of high-performance computing thanks to their massive computational capabilities. Internally, thousands of functional units, architected to be fed by large register files, fuel such a performance. At deep nanometer technologies, the SRAM memory cells that implement GPU register files are very sensitive to the Negative Bias Temperature Instability (NBTI) effect. NBTI ages cell transistors by degrading their threshold voltage$V_{th}$over the lifetime of the GPU. This degradation, which manifests when a cell keeps the same logic value for a relatively long period of time, compromises the cell read stability and increases the transistor switching delay, which can lead to wrong read values and eventually exceed the processor cycle time, respectively, so resulting in faulty operation. This work proposes architectural mechanisms leveraging the redundancy of the data stored in GPU register files to attack NBTI aging. The proposed mechanisms are based on data compression, power gating, and register address rotation techniques. All these mechanisms working together balance the distribution of logic values stored in the cells along the execution time, reducing both the overall$V_{th}$degradation and the increase in the transistor switching delays. Experimental results show that a conventional GPU register file suffers the worst case for NBTI, since a significant fraction of the cells maintain the same logic value during the entire application execution (i.e., a 100 percent ‘0’ and ‘1’ duty cycle distributions). On average, the proposal reduces these distributions by 58 and 68 percent, respectively, which translates into$V_{th}$degradation savings by 54 and 62 percent, respectively. Alejandro Valero, Francisco Candel, Darío Suárez Gracia, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 5 |
| 2019 | FOS: a low-power cache organization for multicores
José Puche, Salvador Petit, Julio Sahuquillo, María Engracia Gómez |
J. Supercomput. | 3 |
| 2019 | Way Combination for an Adaptive and Scalable Coherence DirectoryabstractThis manuscript opens the way to a new class of coherence directory structures that are based on the brand-new concept of way combining. A Way-Combining Directory (WC-dir) builds on a typical sparse directory but allows to take advantage of several ways in the same set to codify the sharing information of each memory block. The result is a sparse directory with variable effective associativity per set and variable length entries, thus being able to dynamically adapt the directory structure to the particular requirements of each application. In particular, our proposal uses just enough bits per entry to store a single pointer, which is optimal for the common case of having just one sharer. For those addresses that have more than one sharer, we have observed that in the majority of cases extra bits could be taken from other empty ways in the same set. All in all, our proposal minimizes the storage overheads without losing the flexibility to adapt to several sharing degrees and without the complexities of other previously proposed techniques. Detailed simulations of a 128-core multicore architecture running benchmarks from PARSEC-3.0 and SPLASH-3 demonstrate that WC-dir can closely approach the performance of a non-scalable bit vector sparse directory, beating the state-of-the-art Scalable Coherence Directory (SCD) and Pool directory proposals. J. Rubén Titos Gil, Antonio Flores, Ricardo Fernández-Pascual, Alberto Ros 0001, Salvador Petit, Julio Sahuquillo, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | Improving GPU Cache Hierarchy Performance with a Fetch and Replacement Cache
Francisco Candel, Salvador Petit, Alejandro Valero, Julio Sahuquillo |
Euro-Par | 4 |
| 2018 | Improving System Turnaround Time with Intel CAT by Identifying LLC Critical Applications
Lucia Pons, Vicent Selfa, Julio Sahuquillo, Salvador Petit, Julio Pons |
Euro-Par | 3 |
| 2018 | Accurately modeling the on-chip and off-chip GPU memory subsystem
Francisco Candel, Salvador Petit, Julio Sahuquillo, José Duato |
Future Gener. Comput. Syst. | 3 |
| 2018 | Designing lab sessions focusing on real processors for computer architecture courses: A practical perspective
Josué Feliu, Julio Sahuquillo, Salvador Petit |
J. Parallel Distributed Comput. | 2 |
| 2018 | Efficient selective multicore prefetching under limited memory bandwidth
Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena |
J. Parallel Distributed Comput. | 2 |
| 2017 | Application Clustering Policies to Address System Fairness with Intel's Cache Allocation TechnologyabstractAchieving system fairness is a major design concern in current multicore processors. Unfairness arises due to contention in the shared resources of the system, such as the LLC and main memory. To address this problem, many research works have proposed novel cache partitioning policies aimed at addressing system fairness without harming performance. Unfortunately, existing proposals targeting fairness require extra hardware which makes them impractical in commercial processors.Recent Intel Xeon processors feature Cache Allocation Technology (CAT), a hardware cache partitioning mechanism that can be controlled from userspace software and that allows to create partitions in the LLC and assign different groups of applications to them.In this paper we propose a family of clustering-based cache partitioning policies to address fairness in systems that feature Intel's CAT. The proposal acts at two levels: applications showing similar amount of core stalls due to LLC accesses are first grouped into clusters, after which each cluster is given a number of ways using a simple mathematical model. To the best of our knowledge, this is the first attempt to address system fairness using the cache partitioning hardware in a real product. Results show that our best performing policy reduces system unfairness by up to 80% (39% on average) for 8-application workloads and by up to 45% (25% on average) for 12-application workloads compared to a non-partitioning approach. Vicent Selfa, Julio Sahuquillo, Lieven Eeckhout, Salvador Petit, María Engracia Gómez |
PACT | 2 |
| 2017 | Exploiting Data Compression to Mitigate Aging in GPU Register FilesabstractNowadays, GPUs sit at the forefront of highperformance computing thanks to their massive computational capabilities. Internally, thousands of functional units, architected to be fed by large register files, fuel such a performance.At nanometer technologies, the SRAM cells that implement register files suffer the Negative Bias Temperature Instability (NBTI) effect, which degrades the transistor threshold voltage Vth and, in turn, can make cells faulty unreliable when they hold the same logic value for long periods of time.Fortunately, the GPU single-thread multiple-data execution model writes data in recognizable patterns. This work proposes mechanisms to detect those patterns, and to compress and shuffle the data, so that compressed register file entries can be safely powered off, mitigating NBTI aging.Experimental results show that a conventional GPU register file experiences the worst case for NBTI, since maintains cells with a single logic value during the entire application execution (i.e., a 100% 0 and 1 duty cycle distributions). On average, the proposal reduces these distributions by 61% and 72%, respectively, which translates into Vth degradation savings by 57% and 64%, respectively. Francisco Candel, Alejandro Valero, Salvador Petit, Darío Suárez Gracia, Julio Sahuquillo |
SBAC-PAD | 5 |
| 2017 | A research-oriented course on Advanced Multicore Architecture: Contents and active learning methodologies
Salvador Petit, Julio Sahuquillo, María Engracia Gómez, Vicent Selfa |
J. Parallel Distributed Comput. | 2 |
| 2017 | The Tag Filter Architecture: An energy-efficient cache and directory design
Joan J. Valls, Alberto Ros 0001, María Engracia Gómez, Julio Sahuquillo |
J. Parallel Distributed Comput. | 4 |
| 2017 | Perf&Fair: A Progress-Aware Scheduler to Enhance Performance and Fairness in SMT MulticoresabstractNowadays, high performance multicore processors implement multithreading capabilities. The processes running concurrently on these processors are continuously competing for the shared resources, not only among cores, but also within the core. While resource sharing increases the resource utilization, the interference among processes accessing the shared resources can strongly affect the performance of individual processes and its predictability. In this scenario, process scheduling plays a key role to deal with performance and fairness. In this work we present a process scheduler for SMT multicores that simultaneously addresses both performance and fairness. This is a major design issue since scheduling for only one of the two targets tends to damage the other. To address performance, the scheduler tackles bandwidth contention at the L1 cache and main memory. To deal with fairness, the scheduler estimates the progress experienced by the processes, and gives priority to the processes with lower accumulated progress. Experimental results on an Intel Xeon E5645 featuring six dual-threaded SMT cores show that the proposed scheduler improves both performance and fairness over two state-of-the-art schedulers and the Linux OS scheduler. Compared to Linux, unfairness is reduced to a half while still improving performance by 5.6 percent. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IEEE Trans. Computers | 2 |
| 2017 | Improving IBM POWER8 Performance Through Symbiotic Job SchedulingabstractSymbiotic job scheduling, i.e., scheduling applications that co-run well together on a core, can have a considerable impact on the performance of processors with simultaneous multithreading (SMT) cores. SMTcores share most of their microarchitectural components among the co-running applications, which causes performance interference between them. Therefore, scheduling applications with complementary resource requirements on the same core can greatly improve the throughput of the system. This paper enhances symbiotic job scheduling for the IBM POWER8 processor. We leverage the existing cycle accounting mechanism to build an interference model that predicts symbiosis between applications. The proposed models achieve higher accuracy than previous models by predicting job symbiosis from throttled CPI stacks, i.e., CPI stacks of the applications when running in the same SMT mode to consider the statically partitioned resources, but without interference from other applications. The symbiotic scheduler uses these interference models to decide, at run-time, which applications should run on the same core or on separate cores. We prototype the symbiotic scheduler as a user-level scheduler in the Linux operating system and evaluate it on an IBM POWER8 server running multiprogram workloads. The symbiotic job scheduler significantly improves performance compared to both an agnostic random scheduler and the default Linux scheduler. Across all evaluated workloads in SMT4 mode, throughput improves by 12.4 and 5.1 percent on average over the random and Linux schedulers, respectively. Josué Feliu, Stijn Eyerman, Julio Sahuquillo, Salvador Petit, Lieven Eeckhout |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared CachesabstractShared caches have become the common design choice in the vast majority of modern multi-core and many-core processors, since cache sharing improves throughput for a given silicon area. Sharing the cache, however, has a downside: the requests from multiple applications compete among them for cache resources, so the execution time of each application increases over isolated execution. The degree in which the performance of each application is affected by the interference becomes unpredictable yielding the system to unfairness situations. This paper proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead hardware-based cache partitioning approach that addresses system fairness. FPCP reduces the interference by allocating to each application a cache partition and adjusting the partition sizes at runtime. To adjust partitions, our approach estimates during multicore execution the time each application would have taken in isolation, which is challenging. The proposed approach has two main differences over existing approaches. First, FPCP distributes cache ways incrementally, which makes the proposal less prone to estimation errors. Second, the proposed algorithm is much less costly than the state-of-the-art ASM-Cache approach. Experimental results show that, compared to ASM-Cache, FPCP reduces unfairness by 48 percent in four-application workloads and by 28 percent in eight-application workloads, without harming the performance. Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | On Microarchitectural Mechanisms for Cache Wearout ReductionabstractHot carrier injection (HCI) and bias temperature instability (BTI) are two of the main deleterious effects that increase a transistor's threshold voltage over the lifetime of a microprocessor. This voltage degradation causes slower transistor switching and eventually can result in faulty operation. HCI manifests itself when transistors switch from logic “0” to “1” and vice versa, whereas BTI is the result of a transistor maintaining the same logic value for an extended period of time. These failure mechanisms are especially acute in those transistors used to implement the SRAM cells of first-level (L1) caches, which are frequently accessed, so they are critical to performance, and they are continuously aging. This paper focuses on microarchitectural solutions to reduce transistor aging effects induced by both HCI and BTI in the data array of L1 data caches. First, we show that the majority of cell flips are concentrated in a small number of specific bits within each data word. In addition, we also build upon the previous studies, showing that logic “0” is the most frequently written value in a cache by identifying which cells hold a given logic value for a significant amount of time. Based on these observations, this paper introduces a number of architectural techniques that spread the number of flips evenly across memory cells and reduce the amount of time that logic “0” values are stored in the cells by switching OFF specific data bytes. Experimental results show that the threshold voltage degradation savings range from 21.8% to 44.3% depending on the application. Alejandro Valero, Negar Miralaei, Salvador Petit, Julio Sahuquillo, Timothy M. Jones 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Student Research Poster: A Low Complexity Cache Sharing Mechanism to Address System FairnessabstractShared caches have become, de facto, the common design choice in current multi-cores, ranging from embedded devices to high-performance processors. In these systems, requests from multiple applications compete for the cache resources, degrading to different extents their progress, quantified as the performance of individual applications compared to isolated execution. The difference between the progresses of the running applications yields the system to unpredictable behavior and causes a fairness problem. This problem can be addressed by carefully partitioning cache resources among the contending applications, but to be effective, a partitioning approach needs to estimate per-application progress. This work proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead cache partitioning approach which addresses fairness by distributing cache resources among applications depending on their progress. To estimate progress, we have implemented two state-of-the-art performance models, ASM and PTCA, which estimate, at runtime, the performance a given application would have if executed in isolation. Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez |
PACT | 2 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 22 |
| 2016 | A Directory Cache with Dynamic Private-Shared PartitioningabstractAs the core counts increase in each chip multiprocessor generation, coherence protocols should improve scalability inperformance, area, and energy consumption to meet the demandsof larger core counts. Directory-based protocols constitute themost scalable alternative. A conventional directory, however, suffers from an inefficient use of storage and energy. First, thelarge, non-scalable, sharer vectors consume unnecessary area andleakage, especially considering that most of the blocks trackedin a directory are cached by a single core. Second, althoughincreasing directory size and associativity could boost systemperformance, it would come at expenses of energy consumption. This paper proposes the Dynamic Way Partitioning (DWP) Directory, a directory structure that exploits three main workloadcharacteristics to achieve area and energy reductions. First, it iswidely known that even in parallel workloads most of the accessedcache blocks are private. Second, most directory accesses targetthe small number of shared blocks. Third, the shared/privateratio of entries in the directory varies across applications andacross different execution phases within the applications. To takeadvantage of these three characteristics, DWP-Directory reducesthe number of ways with storage for shared blocks and it allowsthis storage to be powered off or on at run-time according to thedynamic requirements of the applications. DWP-Directory is compared to a conventional directory cachewith different associativity degrees and with two state-of-the-artschemes: PS-Directory and Hybrid Representation. Experimentalresults for 32-core CMPs show that DWP-Directory achievesthe best of both worlds: similar performance as a traditionaldirectory with high associativity, and similar area as recentstate-of-the-art schemes. In addition, DWP-Directory reducesstatic and dynamic power consumption by 38.0% and 67.4%,respectively compared to conventional sparse directories. Joan J. Valls, María Engracia Gómez, Alberto Ros 0001, Julio Sahuquillo |
HiPC | 4 |
| 2016 | Symbiotic job scheduling on the IBM POWER8abstractSimultaneous multithreading (SMT) processors share most of the microarchitectural core components among the co-running applications. The competition for shared resources causes performance interference between applications. Therefore, the performance benefits of SMT processors heavily depend on the complementarity of the co-running applications. Symbiotic job scheduling, i.e., scheduling applications that co-run well together on a core, can have a considerable impact on the performance of a processor with SMT cores. Prior work uses sampling or novel hardware support to perform symbiotic job scheduling, which has either a non-negligible overhead or is impossible to use on existing hardware. This paper proposes a symbiotic job scheduler for the IBM POWER8 processor. We leverage the existing cycle accounting mechanism to predict symbiosis between applications, and use that information at run-time to decide which applications should run on the same core or on separate cores. We implement the scheduler in the Linux operating system and evaluate it on an IBM POWER8 server running multiprogrammed workloads. The symbiotic job scheduler significantly improves performance compared to both an agnostic random scheduler and the default Linux scheduler. With respect to Linux, it achieves an average speedup by 8.8% for workloads comprising 12 applications, and by 4.7% on average across all evaluated workloads. Josué Feliu, Stijn Eyerman, Julio Sahuquillo, Salvador Petit |
HPCA | 3 |
| 2016 | Impact of Memory-Level Parallelism on the Performance of GPU Coherence ProtocolsabstractGraphics Processing Units (GPUs) are being implemented in heterogeneous CPU/GPU systems due their high efficiency when executing massively parallel applications. New challenges appear to deal with heterogenous coherence in these systems due to the huge amount (hundreds or thousands) of on-going memory requests of GPUs, which is limited by the Miss Status Holding Register (MSHR) file size associated to the L1 cache. This paper analyzes how the number of MSHRs i) affects to typical memory performance metrics and ii) impacts on the system performance under two recent GPU coherence protocols, called NMOESI and SI (Southern Islands), which introduce distinct coherence traffic. We find two key findings that can help improve the performance of coherence protocols. First, there is a strong correlation between system performance and memory subsystem latency regardless of the used protocol. Second, system performance varies with the number of supported cache misses, however, counterintuitively, supporting more cache misses does not always bring enhanced performance but it can turn into performance drops. Francisco Candel, Salvador Petit, Julio Sahuquillo, José Duato |
PDP | 3 |
| 2016 | A Simple Activation/Deactivation Prefetching Scheme for Chip MultiprocessorsabstractPrefetching significantly reduces the memory latencies of a wide range of applications and thus increases the system performance. However, as a speculative technique, prefetching may also noticeably increase the number of memory accesses, which in turns may negatively impact on the main memory bandwidth consumption, performance, and power. Main memory bandwidth consumption is a critical resource especially in the context of current multicore processors since memory requests from all the cores, both prefetch and demand requests, compete among them in the access to the DRAM banks. Consequently, demand requests may be delayed hurting the system performance. This work proposes the Activation/Deactivation Policies (ADP) scheme for hardware prefetchers in multicore processors. This scheme relies on activation policies that turn on the prefetcher on a given core when it is expected that prefetches will improve the performance, and turn off the prefetcher of that core when it is foreseen that performance will be scarcely improved or not improved at all. The proposed mechanism effectively reduces the memory bandwidth requirements of some cores with respect to a typical always prefetching mechanism, so making available extra bandwidth to the co-runners. Results in a four-core processor show that ADP prefetching achieves similar performance ±2.5% as always prefetching, while significantly reducing the memory bandwidth consumed by use-less prefetches. Moreover, in some applications this reduction is as much as 50%. ADP prefetching is applicable to stream-based prefetchers, global-history-buffer delta correlation prefetchers, and PC-based stride prefetchers. Vicent Selfa, Crispín Gómez Requena, María Engracia Gómez, Julio Sahuquillo |
PDP | 4 |
| 2016 | A dynamic execution time estimation model to save energy in heterogeneous multicores running periodic tasks
Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Luis March, José Duato |
Future Gener. Comput. Syst. | 1 |
| 2016 | Bandwidth-Aware On-Line Scheduling in SMT MulticoresabstractThe memory hierarchy plays a critical role on the performance of current chip multiprocessors. Main memory is shared by all the running processes, which can cause important bandwidth contention. In addition, when the processor implements SMT cores, the L1 bandwidth becomes shared among the threads running on each core. In such a case, bandwidth-aware schedulers emerge as an interesting approach to mitigate the contention. This work investigates the performance degradation that the processes suffer due to memory bandwidth constraints. Experiments show that main memory and L1 bandwidth contention negatively impact the process performance; in both cases, performance degradation can grow up to 40 percent for some of applications. To deal with contention, we devise a scheduling algorithm that consists of two policies guided by the bandwidth consumption gathered at runtime. The process selection policy balances the number of memory requests over the execution time to address main memory bandwidth contention. The process allocation policy tackles L1 bandwidth contention by balancing the L1 accesses among the L1 caches. The proposal is evaluated on a Xeon E5645 platform using a wide set of multiprogrammed workloads, achieving performance benefits up to 6.7 percent with respect to the Linux scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IEEE Trans. Computers | 2 |
| 2015 | Addressing Fairness in SMT Multicores with a Progress-Aware SchedulerabstractCurrent SMT (simultaneous multithreading) processors co-schedule jobs on the same core, thus sharing core resources like L1 caches. In SMT multicores, threads also compete among themselves for uncore resources like the LLC (last level cache) and DRAM modules. Per process performance degradation over isolated execution mainly depends on process resource requirements and the resource contention induced by co-runners. Consequently, the running processes progress at different pace. If schedulers are not progress aware, the unpredictable execution time caused by unfairness can introduce undesirable behaviors on the system such as difficulties to keep priority-based scheduling. This work proposes a job scheduler for SMT multicores that provides fairness to the execution of multi programmed workloads. To this end, the scheduler estimates per-process standalone performance by periodically creating low-contention co-schedules. These estimates are used to compute the per process progress. Then, those processes with less progress are prioritized to enhance fairness. Experimental results on a Intel Xeon with six dual-threaded SMT cores show that the proposed scheduler reduces unfairness, on average, by 3× over Linux OS. Moreover, thanks to the tread to core allocation policy, the scheduler slightly improves throughput and turnaround time. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IPDPS | 2 |
| 2015 | Row Tables: Design Choices to Exploit Bank Locality in Multiprogram WorkloadsabstractMain memory is a major performance bottleneck in current chip multiprocessors. Current DRAM banks latch the last accessed row in an internal buffer, namely row buffer (RB), which allows fast subsequent accesses to that row. This throughput-oriented approach was originally designed for single-thread processors and pursues to take advantage of the spatial locality that individual applications exhibit. This paper proposes row tables, a pool of row buffers shared among threads. Depending on the needs of each thread, row buffers are dynamically allocated to threads. Two design approaches are devised differing on the table location, and referred to as BRT (Bank Row Table) and CRT (Controller Row Table), which place the table at the bank, as traditionally done in existing modules, and at the memory controller side, respectively. CRT performs better than BRT in high RB locality applications (or mixes) but performs worse in poor RB locality applications since the increase in transfer times is not later amortized. A variant of CRT referred to as CRT1/xhas been devised to reduce this performance penalty. Results for a 4-core system show that, on average, BRT and CRT1/xmechanisms save energy by 23% and 7%-16% (depending on the X value) and improve IPC by 10% and 9%-14%, respectively. Paula Navarro, Vicent Selfa, Julio Sahuquillo, María Engracia Gómez, Crispín Gómez Requena |
PDP | 3 |
| 2015 | Methodologies and Performance Metrics to Evaluate Multiprogram WorkloadsabstractMulticore processors are dominating the microprocessor market and most research work has moved to this kind of processors. Multicore research methods are still immature and evolving from the single-threaded processor ounterparts. Three main research issues must be faced when evaluating performance and energy in multicores. First, multiple simulation methodologies are being applied to evaluate these systems, without being an agreement about which to use. Second, due to the nature of multiprogram workloads new performance metrics are required, different from those used in single-thread processors. Many metrics have been defined and distinct metrics are used across the published works. Finally, multicore processors are really complex systems which require from sophisticated and complementary (e.g. energy and performance) simulators. This paper pursues to help researchers face the three mentioned research issues. For this purpose, we compare these issues across 28 papers published in 2013 in top computer architecture conferences. Both analytical examples and experimental results are presented with the aim of providing some insights in multicore research. Vicent Selfa, Julio Sahuquillo, Crispín Gómez Requena, María Engracia Gómez |
PDP | 2 |
| 2015 | The Tag Filter Cache: An Energy-Efficient ApproachabstractPower consumption in current high-performance chip multiprocessors (CMPs) has become a major design concern. The current trend of increasing the core count aggravates this problem. On-chip caches consume a significant fraction of the total power budget. Most of the proposed techniques to reduce the energy consumption of these memory structures are at the cost of performance, which may become unacceptable for high-performance CMPs. On-chip caches in multi-core systems are usually deployed with a high associativity degree in order to enhance performance. Even first-level caches are currently implemented with eight ways. The concurrent access to all the ways in the cache set is costly in terms of energy. In this paper we propose an energy-efficient cache design, namely the Tag Filter Cache (TF-Cache) architecture, that filters some of the set ways during cache accesses, allowing to access only a subset of them without hurting the performance. Our cache for each way stores the lowest order tag bits in an auxiliary bit array and these bits are used to filter the ways that do not match those bits in the searched block tag. Experimental results show that, on average, the TF-Cache architecture reduces the dynamic power consumption up to 74.9% and 85.9% when applied to the L1 and L2 cache, respectively, for the studied applications. Joan J. Valls, Julio Sahuquillo, Alberto Ros 0001, María Engracia Gómez |
PDP | 2 |
| 2015 | Surfing the Web Using Browser Interface Facilities: A Performance Evaluation Approach
Raúl Peña-Ortiz, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
J. Web Eng. | 3 |
| 2015 | Design of Hybrid Second-Level CachesabstractIn recent years, embedded dynamic random-access memory (eDRAM) technology has been implemented in last-level caches due to its low leakage energy consumption and high density. However, the fact that eDRAM presents slower access time than static RAM (SRAM) technology has prevented its inclusion in higher levels of the cache hierarchy. This paper proposes to mingle SRAM and eDRAM banks within the data array of second-level (L2) caches. The main goal is to achieve the best trade-off among performance, energy, and area. To this end, two main directions have been followed. First, this paper explores the optimal percentage of banks for each technology. Second, the cache controller is redesigned to deal with performance and energy. Performance is addressed by keeping the most likely accessed blocks in fast SRAM banks. In addition, energy savings are further enhanced by avoiding unnecessary destructive reads of eDRAM blocks. Experimental results show that, compared to a conventional SRAM L2 cache, a hybrid approach requiring similar or even lower area speedups the performance on average by 5.9 percent, while the total energy savings are by 32 percent. For a 45 nm technology node, the energy-delay-area product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless of the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is an eighth of the total number of cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 2 |
| 2015 | PS-Cache: an energy-efficient cache design for chip multiprocessors
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez |
J. Supercomput. | 3 |
| 2015 | PS directory: a scalable multilevel directory cache for CMPs
Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez |
J. Supercomput. | 3 |
| 2014 | Addressing bandwidth contention in SMT multicores through schedulingabstractTo mitigate the impact of bandwidth contention, which in some processes can yield to performance degradations up to 40%, we devise a scheduling algorithm that tackles main memory and L1 bandwidth contention. Experimental evaluation on a real system shows that the proposal achieves an average speedup by 5% with respect to Linux. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
ICS | 2 |
| 2014 | Cache-Hierarchy Contention-Aware Scheduling in CMPsabstractTo improve chip multiprocessor (CMP) performance, recent research has focused on scheduling strategies to mitigate main memory bandwidth contention. Nowadays, commercial CMPs implement multilevel cache hierarchies that are shared by several multithreaded cores. In this microprocessor design, contention points may appear along the whole memory hierarchy. Moreover, this problem is expected to aggravate in future technologies, since the number of cores and hardware threads, and consequently the size of the shared caches increase with each microprocessor generation. This paper characterizes the impact on performance of the different contention points that appear along the memory subsystem. The analysis shows that some benchmarks are more sensitive to contention in higher levels of the memory hierarchy (e.g., shared L2) than to main memory contention. In this paper, we propose two generic scheduling strategies for CMPs. The first strategy takes into account the available bandwidth at each level of the cache hierarchy. The strategy selects the processes to be coscheduled and allocates them to cores to minimize contention effects. The second strategy also considers the performance degradation each process suffers due to contention-aware scheduling. Both proposals have been implemented and evaluated in a commercial single-threaded quad-core processor with a relatively small two-level cache hierarchy. The proposals reach, on average, a performance improvement by 5.38 and 6.64 percent when compared with the Linux scheduler, while this improvement is by 3.61 percent for an state-of-the-art memory contention-aware scheduler under the evaluated mixes. Josué Feliu, Salvador Petit, Julio Sahuquillo, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Efficient Register Renaming and Recovery for High-Performance ProcessorsabstractModern superscalar processors implement register renaming using either random access memory (RAM) or content-addressable memories (CAM) tables. The design of these structures should address both access time and misprediction recovery penalty. Although direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. The presence of associative ports in CAMs, however, prevents them from scaling with the number of physical registers and pipeline width, negatively impacting performance, area, and energy consumption at the rename stage. In this paper, we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides fast and energy-efficient access to register mappings. On misspeculation, a low-complexity CAM enables immediate recovery. Experimental results show that in a four-way state-of-the-art superscalar processor, the new approach provides almost the same performance as an ideal CAM-based renaming scheme, while dissipating only between 17% and 26% of the original energy and, in some cases, consuming less energy than purely RAM-based renaming schemes. Overall, the silicon area required to implement the hybrid RAM-CAM scheme does not exceed the area required by conventional renaming mechanisms. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | L1-bandwidth aware thread allocation in multicore SMT processorsabstractImproving the utilization of shared resources is a key issue to increase performance in SMT processors. Recent work has focused on resource sharing policies to enhance the processor performance, but their proposals mainly concentrate on novel hardware mechanisms that adapt to the dynamic resource requirements of the running threads. This work addresses the L1 cache bandwidth problem in SMT processors experimentally on real hardware. Unlike previous work, this paper concentrates on thread allocation, by selecting the proper pair of co-runners to be launched to the same core. The relation between L1 bandwidth requirements of each benchmark and its performance (IPC) is analyzed. We found that for individual benchmarks, performance is strongly connected to L1 bandwidth consumption, and this observation remains valid when several co-runners are launched to the same SMT core. Based on these findings we propose two L1 bandwidth aware thread to core (t2c) allocation policies, namely Static and Dynamic t2c allocation, respectively. The aim of these policies is to properly balance L1 bandwidth requirements of the running threads among the processor cores. Experiments on a Xeon E5645 processor show that the proposed policies significantly improve the performance of the Linux OS kernel regardless the number of cores considered. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
PACT | 2 |
| 2013 | PS-cache: An energy-efficient cache design for chip multiprocessorsabstractAs silicon resources become increasingly abundant, core counts grow rapidly in successive chip-multiprocessors (CMP) generations. Parallel workloads represent an important segment for current and future CMPs mainly when many-core processors are considered. Unlike multiprogrammed workloads, the accessed blocks in these workloads can be classified in two categories: private, accessed only by one core, and shared, accessed by several cores. This paper takes advantage of this classification to access only a subset of the ways on each L1 cache access, thus reducing dynamic power consumption. Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez |
PACT | 3 |
| 2013 | Combining RAM technologies for hard-error recovery in L1 data caches working at very-low power modesabstractLow-power modes in modern microprocessors rely on low frequencies and low voltages to reduce the energy budget. Nevertheless, manufacturing induced parameter variations can make SRAM cells unreliable producing hard errors at supply voltages below Vccmin. Vicente Lorente, Alejandro Valero, Julio Sahuquillo, Salvador Petit, Ramon Canal, Pedro López 0001, José Duato |
DATE | 3 |
| 2013 | Exploiting reuse information to reduce refresh energy in on-chip eDRAM cachesabstractThis work introduces a novel refresh mechanism that leverages reuse information to decide which blocks should be refreshed in an energy-aware eDRAM last-level cache. Experimental results show that, compared to a conventional eDRAM cache, the energy-aware approach achieves refresh energy savings up to 71%, while the reduction on the overall dynamic energy is by 65% with negligible performance losses. Alejandro Valero, Julio Sahuquillo, Salvador Petit, José Duato |
ICS | 2 |
| 2013 | Referrer Graph: A cost-effective algorithm and pruning method for predicting web accesses
Bernardo de la Ossa, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
Comput. Commun. | 3 |
| 2013 | Analyzing web server performance under dynamic user workloads
Raúl Peña-Ortiz, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
Comput. Commun. | 3 |
| 2013 | Power-aware scheduling with effective task migration for real-time multicore embedded systemsabstractSUMMARY A major design issue in embedded systems is reducing the power consumption because batteries have a limited energy budget. For this purpose, several techniques such as dynamic voltage and frequency scaling (DVFS) or task migration are being used. DVFS allows reducing power by selecting the optimal voltage supply, whereas task migration achieves this effect by balancing the workload among cores. This paper focuses on power‐aware scheduling allowing task migration to reduce energy consumption in multicore embedded systems implementing DVFS capabilities. To address energy savings, the devised schedulers follow two main rules: migrations are allowed at specific points of time and only one task is allowed to migrate each time. Two algorithms have been proposed working under real‐time constraints. The simpler algorithm, namely, single option migration (SOM) only checks just one target core before performing a migration. In contrast, the multiple option migration (MOM) searches the optimal target core. In general, the MOM algorithm achieves better energy savings than the SOM algorithm, although differences are wider for a reduced number of cores and frequency/voltage levels. Moreover, the MOM algorithm reduces energy consumption as much as 40% over the worst fit algorithm. Copyright © 2012 John Wiley & Sons, Ltd. José Luis March, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Hardware-Based Generation of Independent Subtraces of Instructions in Clustered ProcessorsabstractMulticore chips are currently dominating the microprocessor market as designs that improve performance and sustain power consumption. However, complex core features must be still considered to provide good performance for existing sequential applications. An effective approach to reduce core complexity without dramatically sacrificing performance is to distribute critical processor structures by using clustered microarchitectures. In these designs, communication latency among clusters is a critical performance bottleneck, and a good steering algorithm is required to reduce intercluster communication. In this paper, we propose a new energy-efficient microarchitectural approach that reduces intercluster communication by detecting and generating independent chains of instructions, referred to as subtraces, from the execution of sequential programs. The devised mechanism has been modeled on an x86-based trace-cache processor, where subtraces are built in the fill unit, stored in a trace cache, and individually steered to different clusters. Experimental results show that the proposal reaches performance speedups around 7 and 15 percent for point-to-point and bus-based interconnects, respectively, while achieving energy savings of up to 12 percent. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 2 |
| 2012 | PS-Dir: a scalable two-level directory cacheabstractAs the number of cores increases in both incoming and future chip multiprocessors, coherence protocols must address novel hardware structures in order to scale in terms of performance, power, and area. It is well known that most blocks accessed by parallel applications are private (i.e., accessed by a single core). These blocks present different directory requirements and behavior than shared blocks. Based on this fact, this paper proposes a two-level directory cache that tracks shared blocks in a small and fast first-level cache and private blocks in a larger and slower second-level cache, namely Shared and Private caches, respectively. Speed and area reasons suggest the use of eDRAM technology much dense but slower than SRAM technology for the Private cache, which in turn brings energy savings. Experimental results for a 16-core system show improvements in performance by 11.1%, in area by 25.4%, and in energy consumption by 20.5% compared to a conventional directory cache. Joan J. Valls, Alberto Ros 0001, Julio Sahuquillo, María Engracia Gómez, José Duato |
PACT | 3 |
| 2012 | Analyzing the optimal ratio of SRAM banks in hybrid cachesabstractCache memories have been typically implemented with Static Random Access Memory (SRAM) technology. This technology presents a fast access time but high energy consumption and low density. As opposite, the recently appeared embedded Dynamic RAM (eDRAM) technology allows caches to be built with lower energy and area, although with a slower access time. The eDRAM technology provides important leakage and area savings, especially in huge Last-Level Caches (LLCs), which occupy almost half the silicon area in some recent microprocessors. This paper proposes a novel hybrid LLC, which combines SRAM and eDRAM banks to address the trade-off among performance, energy, and area. To this end, we explore the optimal percentage of SRAM and eDRAM banks that achieves the best target trade-off. Architectural mechanisms have been devised to keep the most likely accessed blocks in fast SRAM banks as well as to avoid unnecessary destructive reads. Experimental results show that, compared to a conventional SRAM LLC with the same storage capacity, performance degradation does not surpass, on average, 2.9% (even with 12.5% of banks built with SRAM technology), whereas area savings can be as high as 46% for a 1MB-16way LLC. For a 45nm technology node, the energy-delay squared product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is a quarter or an eighth of the cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ICCD | 2 |
| 2012 | Page-Based Memory Allocation Policies of Local and Remote Memory in Cluster ComputersabstractMain memory latencies have a strong impact on the overall execution time of the applications. The need of efficiently scheduling the costly DRAM memory resources in the different motherboards is a major concern in cluster computers. Most of these systems implement remote access capabilities which allow the OS to access to remote memory. In this context, efficient scheduling becomes even more critical since remote memory accesses may be several orders of magnitude higher than local accesses. These systems typically support interleaved memory at cache-block granularity. In contrast, in this paper we explore the impact on the system performance when allocating memory at OS page granularity. Experimental results show that simply supporting interleaved memory at OS page granularity is a feasible solution that does not impact on the performance of most of the benchmarks. Based on this observation we investigated the reasons of performance drops in those benchmarks showing unacceptable performance when working at page granularity. The results of this analysis lead us to propose two memory allocation policies, namely on-demand (OD) and Most-accessed in-local (Mail). The OD policy first places the requested pages in local memory, once this memory region is full, the subsequent memory pages are placed in remote memory. This policy shows good performance when the most accessed pages are requested and allocated before than the least accessed ones, which as proven in this work, is the most common case. This simple policy reaches performance improvements by 25% in some benchmarks with respect to a typical block interleaving memory system. Nevertheless, this strategy has poor performance when a noticeable amount of the least accessed pages are requested before than the most accessed ones. This performance drawback is solved by the Mail allocation policy by using profile information to guide the allocation of new pages. This scheme always outperforms the baseline block interleaving policy and, in some cases, improves the performance of the OD policy by 25%. Monica Serrano, Salvador Petit, Julio Sahuquillo, Rafael Ubal, Houcine Hassan, José Duato |
ICPADS | 3 |
| 2012 | Understanding Cache Hierarchy Contention in CMPs to Improve Job SchedulingabstractIn order to improve CMP performance, recent research has focused on scheduling to mitigate contention produced by the limited memory bandwidth. Nowadays, commercial CMPs implement multi-level cache hierarchies where last level caches are shared by at least two cache structures located at the immediately lower cache level. In turn, these caches can be shared by several multithreaded cores. In this microprocessor design, contention points may appear along the whole memory hierarchy. Moreover, this problem is expected to aggravate in future technologies, since the number of cores and hardware threads, and consequently the size of the shared caches increases with each microprocessor generation. In this paper we characterize the impact on performance of the different contention points that appear along the memory subsystem. Then, we propose a generic scheduling strategy for CMPs that takes into account the available bandwidth at each level of the cache hierarchy. The proposed strategy selects the processes to be co-scheduled and allocates them to cores in order to minimize contention effects. The proposal has been implemented and evaluated in a commercial single-threaded quad-core processor with a relatively small two-level cache hierarchy. Despite these potential contention limitations are less than in recent processor designs, compared to the Linux scheduler, the proposal reaches performance improvements up to 9% while these benefits (across the studied benchmark mixes) are always lower than 6% for a memory-aware scheduler that does not take into account the cache hierarchy. Moreover, in some cases the proposal doubles the speedup achieved by the memory-aware scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IPDPS | 2 |
| 2012 | The Impact of User's Dynamic Behavior on Web PerformanceabstractThe increasing popularity of web applications has introduced a new paradigm where users are no longer passive web consumers but they become active contributors to the Web, specially in the contexts of social networking, blogs, wikis or e-commerce. In this new paradigm, contents and services are even more dynamic, which consequently increases the level of dynamism in user's behavior. Moreover, this trend is expected to rise in the incoming Web. This dynamism is a major adversity to define and model representative web workload, in fact, this characteristic is not fully represented in the most of the current web workload generators. This work proves that the web user's dynamic behavior is a crucial point that must be addressed in web performance studies in order to accurately estimate system performance indexes. In this paper, we analyze the effect of using a more realistic dynamic workload on the web performance metrics. To this end, we evaluate a typical e-commerce scenario and compare the results obtained using dynamic workload instead of traditional workloads. Experimental results show that, when a more dynamic and interactive workload is taken into account, performance indexes can widely differ and noticeably affect the stress borderline on the server. For instance, the processor usage can increase 20% due to dynamism, affecting negatively average response time perceived by users, which can also turn in unwanted effects in marketing and fidelity policies. Raúl Peña-Ortiz, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
NCA | 3 |
| 2012 | Efficiently Handling Memory Accesses to Improve QoS in Multicore Systems under Real-Time ConstraintsabstractChip multiprocessors (CMPs) are becoming the common choice to implement embedded systems due to they achieve a good tradeoff between performance and power. Because of manufacturability reasons, CMPs use to implement one or several memory controllers, each one shared by a set of cores. Thus, memory requests from distinct cores compete among them when accessing to memory. This means that the memory access latency can widely vary depending on the co-runners and the memory controller scheduling policy, thus yielding to unpredictable behavior. This work focuses on the design of a memory controller to support workloads with real-time constraints, both hard real-time (HRT) and soft real-time (SRT) applications. These systems must guarantee the execution of HRT applications while improving the performance of the SRT applications. In this paper we propose two memory controller policies for multicore embedded systems: HR-first and ATR-first. The former prioritizes memory requests of HRT tasks, achieving important energy savings but poor performance for SRT applications. The latter gives priority to those HRT requests that are critical to guarantee schedulability. Results show that the ATR-first policy presents similar energy consumption as the HR-first policy while reducing the number of SRT deadline misses around 49%, on average, and reaching the fulfillment of all deadlines in some scenarios. José Luis March, Salvador Petit, Julio Sahuquillo, Houcine Hassan, José Duato |
SBAC-PAD | 3 |
| 2012 | A taxonomy of web prediction algorithms
Josep Domenech 0001, Bernardo de la Ossa, Julio Sahuquillo, José A. Gil 0001, Ana Pont |
Expert Syst. Appl. | 3 |
| 2012 | Key factors in web latency savings in an experimental prefetching system
Bernardo de la Ossa, Julio Sahuquillo, Ana Pont, José A. Gil 0001 |
J. Intell. Inf. Syst. | 2 |
| 2012 | Prediction Algorithms for Prefetching in the Current Web
Josep Domenech 0001, Julio Sahuquillo, José A. Gil 0001, Ana Pont |
J. Web Eng. | 2 |
| 2012 | Combining recency of information with selective random and a victim cache in last-level cachesabstractMemory latency has become an important performance bottleneck in current microprocessors. This problem aggravates as the number of cores sharing the same memory controller increases. To palliate this problem, a common solution is to implement cache hierarchies with large or huge Last-Level Cache (LLC) organizations. LLC memories are implemented with a high number of ways (e.g., 16) to reduce conflict misses. Typically, caches have implemented the LRU algorithm to exploit temporal locality, but its performance goes away from the optimal as the number of ways increases. In addition, the implementation of a strict LRU algorithm is costly in terms of area and power. This article focuses on a family of low-cost replacement strategies, whose implementation scales with the number of ways while maintaining the performance. The proposed strategies track the accessing order for just a few blocks, which cannot be replaced. The victim is randomly selected among those blocks exhibiting poor locality. Although, in general, the random policy helps improving the performance, in some applications the scheme fails with respect to the LRU policy leading to performance degradation. This drawback can be overcome by the addition of a small victim cache of the large LLC. Experimental results show that, using the best version of the family without victim cache, MPKI reduction falls in between 10% and 11% compared to a set of the most representative state-of-the-art algorithms, whereas the reduction grows up to 22% with respect to LRU. The proposal with victim cache achieves speedup improvements, on average, by 4% compared to LRU. In addition, it reduces dynamic energy, on average, up to 8%. Finally, compared to the studied algorithms, hardware complexity is largely reduced by the baseline algorithm of the family. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Design, Performance, and Energy Consumption of eDRAM/SRAM Macrocells for L1 Data CachesabstractSRAM and DRAM have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology (eDRAM), thus overcoming the speed limit of typical DRAM cells. In this paper, we propose a hybrid n-bit macrocell that implements one SRAM cell and n-1 eDRAM cells. This cell is aimed at being used in an n-way set-associative first-level data cache. Architectural mechanisms (e.g., special writeback policies) have been devised to completely avoid refresh logic. Performance, energy, and area have been analyzed in detail. Experimental results show that using typical eDRAM capacitors, and compared to a conventional cache, a 4-way set-associative hybrid cache reduces both energy consumption and area up to 54 and 29 percent, respectively, while having negligible impact on performance (less than 2 percent). Alejandro Valero, Salvador Petit, Julio Sahuquillo, Pedro López 0001, José Duato |
IEEE Trans. Computers | 3 |
| 2012 | A cost-effective heuristic to schedule local and remote memory in cluster computers
Monica Serrano, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
J. Supercomput. | 2 |
| 2012 | A Sequentially Consistent Multiprocessor Architecture for Out-of-Order Retirement of InstructionsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Expanding the width of the instruction window can be highly beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory latencies. Based on the Validation Buffer (VB) architecture (a previously proposed out-of-order retirement, checkpoint-free architecture for single processors), this paper proposes a cost-effective, scalable, out-of-order retirement multiprocessor, capable of enforcing sequential consistency without impacting the design of the memory hierarchy or interconnect. Our simulation results indicate that utilizing a VB can speed up both relaxed and sequentially consistent in-order retirement in future multiprocessor systems by between 3 and 20 percent, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Impact on Performance and Energy of the Retention Time and Processor Frequency in L1 Macrocell-Based Data CachesabstractCache memories dissipate an important amount of the energy budget in current microprocessors. This is mainly due to cache cells are typically implemented with six transistors. To tackle this design concern, recent research has focused on the proposal of new cache cells. Ann-bit cache cell, namely macrocell, has been proposed in a previous work. This cell combines SRAM and eDRAM technologies with the aim of reducing energy consumption while maintaining the performance. The capacitance of eDRAM cells impacts on energy consumption and performance since these cells lose their state once the retention time expires. On such a case, data must be fetched from a lower level of the memory hierarchy, so negatively impacting on performance and energy consumption. As opposite, if the capacitance is too high, energy would be wasted without bringing performance benefits. This paper identifies the optimal capacitance for a given processor frequency. To this end, the tradeoff between performance and energy consumption of a macrocell-based cache has been evaluated varying the capacitance and frequency. Experimental results show that, compared to a conventional cache, performance losses are lower than 2% and energy savings are up to 55% for a cache with 10 fF capacitors and frequencies higher than 1 GHz. In addition, using trench capacitors, a 4-bit macrocell reduces by 29% the area of four conventional SRAM cells. Alejandro Valero, Julio Sahuquillo, Vicente Lorente, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Improving Last-Level Cache Performance by Exploiting the Concept of MRU-TourabstractLast-Level Caches (LLCs) implement the LRU algorithm to exploit temporal locality, but its performance is quite far of Belady's optimal algorithm as the number of ways increases. One of the main reasons because of LRU does not reach good performance in LLCs is that this policy forces a block to descend until the bottom of the stack before eviction. Nevertheless, most of the blocks that leave the MRU position are not referenced again before eviction. This work pursues to select candidate blocks to be victimized before reaching the bottom of the stack. To this end, this work defines the number of MRU-Tours (MRUTs) of a block as the number of times that a block enters in the MRU position during its live time. Based on the fact that most of the blocks exhibit a single MRUT, this work presents the family of MRUT-based algorithms aimed at exploiting this block behavior to improve performance. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 2 |
| 2011 | Energy Behaviour of NUCA Caches in CMPsabstractAdvances in technology of semiconductor make nowadays possible to design Chip Multiprocessor Systems equipped with huge on-chip Last Level Caches. Due to the wire delay problem, the use of traditional cache memories with a uniform access time would result in unacceptable response latencies. NUCA (Non Uniform Cache Access) architecture has been proposed as a viable solution to hide the adverse impact of wires delay on performance. Many previous studies have focused on the effectiveness of NUCA architectures, but the study of the energy and power aspects of NUCA caches is still limited. In this work, we present an energy model specifically suited for NUCA-based CMP systems, together with a methodology to employ the model to evaluate the NUCA energy consumption. Moreover, we present a performance and energy dissipation analysis for two 8-core CMP systems with an S-NUCA and a D-NUCA, respectively. Experimental results show that, similarly to the monolithic processor, the static power also dominates the total power budget in the CMP system. Alessandro Bardine, Pierfrancesco Foglia, Francesco Panicucci, Marco Solinas, Julio Sahuquillo |
DSD | 5 |
| 2011 | A Dynamic Power-Aware Partitioner with Task Migration for Multicore Embedded Systems
José Luis March, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
Euro-Par (1) | 2 |
| 2011 | A Cluster Computer Performance Predictor for Memory Scheduling
Monica Serrano, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICA3PP (2) | 2 |
| 2011 | MRU-Tour-based Replacement Algorithms for Last-Level CachesabstractMemory hierarchy design is a major concern in current microprocessors. Many research work focuses on the Last-Level Cache (LLC), which is designed to hide the long miss penalty of accessing to main memory. To reduce both capacity and conflict misses, LLCs are implemented as large memory structures with high associativities. To exploit temporal locality, LRU is the replacement algorithm usually implemented in caches. However, for a high-associative cache, its implementation is costly in terms of area and power consumption. Indeed, LRU is not well suited for the LLC, because as this cache level does not see all memory accesses, it cannot cope with temporal locality. In addition, blocks must descend down to the LRU position of the stack before eviction, even when they are not longer useful. In this paper, we show that most of the blocks are not referenced again once they leave the MRU position. Moreover, the probability of being referenced again does not depend on the location on the LRU stack. Based on these observations, we define the number of MRU-Tours (MRUTs) of a block as the number of times that a block occupies the MRU position while it is stored in the cache, and propose the MRUT replacement algorithm, which selects the block to be replaced among the blocks that show only one MRUT. Variations of this algorithm have been also proposed to exploit both MRUT behavior and recency of information. Experimental results show that, compared to LRU, the proposal reduces the MPKI up to 22%, while IPC is improved by 48%. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
SBAC-PAD | 2 |
| 2011 | Web Workload Generators - A Survey Focusing on user Dynamism Representation
Raúl Peña-Ortiz, Julio Sahuquillo, José A. Gil 0001, Ana Pont |
WEBIST | 2 |
| 2011 | A New Energy-Aware Dynamic Task Set Partitioning Algorithm for Soft and Hard Embedded Real-Time SystemsabstractPower consumption is a major design concern in current embedded systems. To deal with consumption, many systems apply dynamic voltage scaling (DVS) techniques which dynamically change the system speed depending on the workload characteristics. DVS costs in a multicore system can be reduced by sharing the same DVS regulator among the cores. In this context, to handle energy efficiently, the workload must be properly balanced among the cores. This paper proposes a new heuristic algorithm to balance the workload in an embedded system with a coarse-grain multithreaded multicore processor. This heuristic is aimed at improving the overlapping time between the memory and the processor while keeping balanced core utilizations. To this end, the heuristic dynamically drives the frequency/voltage level to guarantee deadline fulfillment of the hard real-time tasks as well as to achieve a good trade-off between deadline losses and energy savings of the soft real-time tasks. The proposed technique has been evaluated on a model of a contemporary high-end ARM embedded microprocessor executing a set of standard embedded benchmarks. Energy savings depend on the range of frequency/voltage levels that the DVS regulator implements. Experimental results show that with the proposed heuristic, when working with hard real-time tasks, the energy consumption is about 33% the energy dissipated by a system without DVS regulator and balancing heuristic. Moreover, when soft real-time tasks are also considered, the normalized consumption presents values ranging in between 8 and 70% depending on the scheduler aggressiveness. José Luis March, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
Comput. J. | 2 |
| 2010 | Exploiting subtrace-level parallelism in clustered processorsabstractNo abstract available. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 2 |
| 2010 | A Scheduling Heuristic to Handle Local and Remote Memory in Cluster ComputersabstractIn cluster computers, RAM memory is spread among the motherboards hosting the running applications. In these systems, it is common to constrain the memory address space of a given processor to the local motherboard. Constraining the system in this way is much cheaper than using a full-fledged shared memory implementation among motherboards. However, in this case, memory usage might widely differ among motherboards depending on the memory requirements of the applications running on each motherboard. In this context, if an application requires a huge quantity of RAM memory, the only feasible solution is to increase the amount of available memory in its local motherboard, even if the remaining ones are underused. Nevertheless, beyond a certain memory size, this memory budget increase becomes prohibitive. In this paper, we assume that the Remote Memory Access hardware used in a Hyper Transport based system allows applications to allocate the required memory from remote motherboards. We also analyze how the distribution of memory accesses among different memory locations (local or remote) impact on performance. Finally, an heuristic is devised to schedule local and remote memory among applications according to their requirements, and considering quality of service constraints. Monica Serrano, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
HPCC | 2 |
| 2010 | Extending a Multicore Multithread Simulator to Model Power-Aware Hard Real-Time Systems
José Luis March, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICA3PP (2) | 2 |
| 2010 | Out-of-order retirement of instructions in sequentially consistent multiprocessorsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Wide instruction windows are very beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory gap. In this paper, the Validation Buffer (VB) multiprocessor architecture is proposed as a cost-effective, checkpoint-free, scalable approach to retire instructions out of program order, while still enforcing sequential consistency, and without impacting the memory hierarchy or interconnect. Experimental results show that utilizing the Validation Buffer can speed up both release and sequentially consistent in-order retirement in future multiprocessor systems by between 3% and 20%, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
ICCD | 2 |
| 2010 | Speculative Validation of Web Objects for Further Reducing the User-Perceived Latency
Josep Domenech 0001, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
Networking | 3 |
| 2010 | Balancing Task Resource Requirements in Embedded Multithreaded Multicore Processors to Reduce Power ConsumptionabstractPower consumption is a major design issue in modern microprocessors. Hence, power reduction techniques, like Dynamic Voltage Scaling (DVS), are being widely implemented. Unfortunately, they impact on the task execution time so difficulting schedulability of hard real-time applications. To deal with this problem, this paper proposes a power-aware scheduler for coarse-grain embedded multicore processors implementing global DVS. To this end, this work presents two heuristics, namely Balanced Memory and Balanced CPU, which distribute the task set among cores focusing on resource utilization. Results show that with respect to a system not implementing DVS, two or five DVS levels achieve energy savings by about 35% or 51%, respectively. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
PDP | 2 |
| 2010 | Using current web page structure to improve prefetching performance
Josep Domenech 0001, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
Comput. Networks | 3 |
| 2009 | An Efficient Low-Complexity Alternative to the ROB for Out-of-Order Retirement of InstructionsabstractCurrent superscalar processors use a reorder buffer (ROB) to support speculation, precise exceptions, and register reclamation. Instructions are retired from this structure in program order, which may lead to significant performance degradation if a long latency operation blocks the ROB head. In this paper, a checkpoint-free out-of-order commit architecture is proposed, which replaces the ROB with a small structure called validation buffer (VB) from which instructions are retired as soon as their speculative state is resolved. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. Experimental results show that the VB microarchitecture is much more efficient than a ROB-based microprocessor. For example, a 32-entry VB provides similar performance to a 256-entry ROB, while reducing the utilization of other major processor structures. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001, José Duato |
DSD | 3 |
| 2009 | Paired ROBs: A Cost-Effective Reorder Buffer Sharing Strategy for SMT Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 2 |
| 2009 | A power-aware hybrid RAM-CAM renaming mechanism for fast recoveryabstractModern superscalar processors implement register renaming by using either RAM or CAM tables. The design of these structures should address their access time and misprediction recovery penalty. While direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. Although they are more complex and slower, CAMs usually match the processor cycle in current designs. However, they do not scale with the number of physical registers and the pipeline width. In this paper we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides the current mappings quickly; on mispeculation, a low-complexity CAM enables immediate recovery and further register renaming. Compared to an ideal CAM in a 4-way state-of-the-art superscalar microprocessor, and for almost the same performance (1% slowdown) and area (95% of the ideal CAM size), the proposed scheme consumes about 90% less dynamic energy. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
ICCD | 3 |
| 2009 | Dynamic task set partitioning based on balancing memory requirements to reduce power consumptionabstractBecause of technology advances power consumption has emerged up as an important design issue in modern high-performance microprocessors. As a consequence, research on reducing power consumption has become a hot research topic. Different ways to reduce power consumption consist on using processors that do not implement the most power-hungry microarchitectural mechanisms, attacking hot spots, or reducing consumption in the larger microprocessor components like the cache. Unlike these works which focus on specific parts of the microprocessor, Dynamic Voltage Scaling (DVS) is a technique which applies on the whole microprocessor die. This technique allows the system to work at different frequency/voltage levels. DVS costs in a multicore system can be reduced by sharing the same DVS regulator among the cores (global DVS). In this context, to handle energy efficiently, the workload must be properly balanced among the cores. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICS | 2 |
| 2009 | An hybrid eDRAM/SRAM macrocell to implement first-level data cachesabstractSRAM and DRAM cells have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology, thus overcoming the speed limit of typical DRAM cells. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Vicente Lorente, Ramon Canal, Pedro López 0001, José Duato |
MICRO | 2 |
| 2009 | An Empirical Study on Maximum Latency Saving in Web PrefetchingabstractThis paper presents an empirical study to investigate the maximum benefits that web users can expect from prefetching techniques in the current web. To this end a perfect web predictor is defined, but unlike previous theoretical studies, this work considers a realistic prefetching architecture using real and representative traces. In this way, the influence of real implementation constraints can be considered. The results obtained show that web prefetching can improve page latency up to 52% in the studied traces. Bernardo de la Ossa, Julio Sahuquillo, Ana Pont, José A. Gil 0001 |
Web Intelligence | 2 |
| 2009 | Dweb model: Representing Web 2.0 dynamism
Raúl Peña-Ortiz, Julio Sahuquillo, Ana Pont, José A. Gil 0001 |
Comput. Commun. | 2 |
| 2009 | A Complexity-Effective Out-of-Order Retirement MicroarchitectureabstractCurrent superscalar processors commit instructions in program order by using a reorder buffer (ROB). The ROB provides support for speculation, precise exceptions, and register reclamation. However, committing instructions in program order may lead to significant performance degradation if a long latency operation blocks the ROB head. Several proposals have been published to deal with this problem. Most of them retire instructions speculatively. However, as speculation may fail, checkpoints are required in order to rollback the processor to a precise state, which requires both extra hardware to manage checkpoints and the enlargement of other major processor structures, which, in turn, might impact the processor cycle. This paper focuses on out-of-order commit in a nonspeculative way, thus, avoiding checkpointing. To this end, we replace the ROB with a validation buffer (VB) structure. This structure keeps dispatched instructions until they are nonspeculative or mispeculated, which allows an early retirement. By doing so, the performance bottleneck is largely alleviated. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. As experimental results show, the VB structure is much more efficient than a typical ROB since, with only 32 entries, it achieves a performance close to an in-order commit microprocessor using a 256-entry ROB. Salvador Petit, Julio Sahuquillo, Pedro López 0001, Rafael Ubal, José Duato |
IEEE Trans. Computers | 2 |
| 2008 | Reducing the Number of Bits in the BTB to Attack the Branch Predictor Hot-Spot
Noel Tomás, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 2 |
| 2008 | A simple power-aware scheduling for multicore systems when running real-time applicationsabstractHigh-performance microprocessors, e.g., multithreaded and multicore processors, are being implemented in embedded real-time systems because of the increasing computational requirements. These complex microprocessors have two major drawbacks when they are used for real-time purposes. First, their complexity difficults the calculation of the WCET (worst case execution time). Second, power consumption requirements are much larger, which is a major concern in these systems. In this paper we propose a novel soft power-aware real-time scheduler for a state-of-the-art multicore multithreaded processor, which implements dynamic voltage scaling techniques. The proposed scheduler reduces the energy consumption while satisfying the constraints of soft real-time applications. Different scheduling alternatives have been evaluated, and experimental results show that using a fair scheduling policy, the proposed algorithm provides, on average, energy savings ranging from 34% to 74%. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
IPDPS | 2 |
| 2008 | The impact of out-of-order commit in coarse-grain, fine-grain and simultaneous multithreaded architecturesabstractMultithreaded processors in their different organizations (simultaneous, coarse grain and fine grain) have been shown as effective architectures to reduce the issue waste. On the other hand, retiring instructions from the pipeline in an out-of-order fashion helps to unclog the ROB when a long latency instruction reaches its head. This further contributes to maintain a higher utilization of the available issue bandwidth. In this paper, we evaluate the impact of retiring instructions out of order on different multithreaded architectures and different instruction fetch policies, using the recently proposed Validation Buffer microarchitecture as baseline out-of-order commit technique. Experimental results show that, for the same performance, out-of-order commit permits to reduce multithread hardware complexity (e.g., fine grain multithreading with a lower number of supported threads). Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IPDPS | 2 |
| 2007 | VB-MT: Design Issues and Performance of the Validation Buffer Microarchitecture for Multithreaded Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 2 |
| 2007 | Delfos: the Oracle to Predict NextWeb User's AccessesabstractDespite the wide and intensive research efforts focused on Web prediction and prefetching techniques aimed to reduce user's perceived latency, few attempts to implement and use them in real environments have been done, mainly due to their complexity and supposed limitations that low user available bandwidths imposed few years ago. Nevertheless, current user bandwidths open a new scenario for prefetching that becomes again an interesting option to improve web performance. This paper presents Delfos, a framework to perform web predictions and prefetching on a real environment that tries to cover the existing gap between research and praxis. Delfos is integrated in the web architecture without modifying the standard HTTP 1.1 protocol, and acts inserting predictions in the web server side, while prefetchs are carried out by the client. In addition, it can be also used as a flexible framework to evaluate and compare existing prefetching techniques and algorithms and to assist in the design of new ones because it provides detailed statistics reports. Bernardo de la Ossa, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
AINA | 3 |
| 2007 | Multi2Sim: A Simulation Framework to Evaluate Multicore-Multithreaded ProcessorsabstractCurrent microprocessors are based in complex designs, integrating different components on a single chip, such as hardware threads, processor cores, memory hierarchy or interconnection networks. The permanent need of evaluating new designs on each of these components motivates the development of tools which simulate the system working as a whole. In this paper, we present the Multi2Sim simulation framework, which models the major components of incoming systems, and is intended to cover the limitations of existing simulators. A set of simulation examples is also included for illustrative purposes. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 2 |
| 2007 | Analysis of Web-Proxy Cache Replacement Algorithms under Steady-state Conditions
Luis Guillermo Cárdenas, Ana Pont, Julio Sahuquillo, José A. Gil 0001 |
WEBIST (1) | 3 |
| 2007 | A user-focused evaluation of web prefetching algorithms
Josep Domenech 0001, Ana Pont, Julio Sahuquillo, José A. Gil 0001 |
Comput. Commun. | 3 |
| 2006 | Cost-Benefit Analysis of Web Prefetching Algorithms from the User's Point of View
Josep Domenech 0001, Ana Pont, Julio Sahuquillo, José A. Gil 0001 |
Networking | 3 |
| 2006 | Applying the zeros switch-off technique to reduce static energy in data cachesabstractZeros switch-off is a leakage energy reduction technique applicable to cache memories. It works at the cache word level by removing the power supply of all or part of its most significant bytes when they store a zero, taking advantage of the high percentage of zero data bits in common programs. Experimental results, obtained by using the SPEC2000 benchmarks suite, show that the average leakage energy savings reach 60.3% with no IPC loss indeed. The proposed technique can be combined with other existing energy reduction techniques, reaching, on average, 67.3% savings with 0.5% IPC losses Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 2 |
| 2006 | The Impact of the Web Prefetching Architecture on the Limits of Reducing User's Perceived LatencyabstractWeb prefetching is a technique that has been researched for years to reduce the latency perceived by users. For this purpose, several Web prefetching architectures have been used, but no comparative study has been performed to identify the best architecture dealing with prefetching. This paper analyzes the impact of the Web prefetching architecture focusing on the limits of reducing the user's perceived latency. To this end, the factors that constrain the predictive power of each architecture are analyzed and these theoretical limits are quantified. Experimental results show that the best element of the Web architecture to locate a single prediction engine is the proxy, whose implementation could reduce the perceived latency up to 67%. Schemes for collaborative predictors located at diverse elements of the Web architecture are also analyzed. These predictors could dramatically reduce the perceived latency, reaching a potential limit of about 97% for a mixed proxy-server collaborative prediction engine Josep Domenech 0001, Julio Sahuquillo, José A. Gil 0001, Ana Pont |
Web Intelligence | 2 |
| 2006 | Web prefetching performance metrics: A survey
Josep Domenech 0001, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
Perform. Evaluation | 3 |
| 2006 | Addressing a workload characterization study to the design of consistency protocols
Salvador Petit, Julio Sahuquillo, Ana Pont, David R. Kaeli |
J. Supercomput. | 2 |
| 2005 | Performance Comparison of a Web Cache Simulation FrameworkabstractPerformance comparison studies are primarily carried out through real systems or simulation environments. Simulation is the most commonly used method to explore new proposals due to both its flexibility and the relatively reduced time taken to obtain performance results. This paper presents a powerful framework to simulate Web proxy cache systems. Our tool provides a comfortable environment to simulate and explore cache management techniques. In order to validate our framework and show how accurate it executes, a performance comparison has been done. We analyzed the details of a commercial proxy cache system and compare its results with those obtained from our simulator using the most commonly replacement algorithm (LRU). For this purpose, the proposed environment was adapted to match the performance of the real proxy cache. Experimental results show that proxy cache hit ratio deviations fall very close to the real system, since then, never exceeds 3.42%. Luis Guillermo Cárdenas, José A. Gil 0001, Josep Domenech 0001, Julio Sahuquillo, Ana Pont |
AINA | 4 |
| 2005 | Emulating Web Cache Replacement Algorithms versus a Real SystemabstractThis paper presents a powerful framework to simulate Web proxy cache systems. Our tool provides a comfortable environment to simulate and explore cache management techniques. It also includes an extension to design and simulate new structures considering several inter-connected caches which are very convenient for our current research projects. Besides a statistics module is enlarged to obtain supplementary performance measures. We compared the results obtained from our framework against a commercial proxy cache system by using several replacement algorithms and input traces. Experimental results show that proxy cache hit ratio deviations fall very close to the real system, since them never exceeds 3.5%. Although the simulation time varies depending on the input trace size and the modeled management technique, in all experiments run time has been by about several hundred times faster than the time the real system takes. Luis Guillermo Cárdenas, José A. Gil 0001, Julio Sahuquillo, Ana Pont |
ISCC | 3 |
| 2005 | Exploring the performance of split data cache schemes on superscalar processors and symmetric multiprocessors
Julio Sahuquillo, Salvador Petit, Ana Pont, Veljko M. Milutinovic |
J. Syst. Archit. | 1 |
| 2005 | On-Chip Interconnects and Instruction Steering Schemes for Clustered MicroarchitecturesabstractClustering is an effective microarchitectural technique for reducing the impact of wire delays, the complexity, and the power requirements of microprocessors. In this work, we investigate the design of on-chip interconnection networks for clustered superscalar microarchitectures. This new class of interconnects has demands and characteristics different from traditional multiprocessor networks. In particular, in a clustered microarchitecture, a low intercluster communication latency is essential for high performance. We propose some point-to-point cluster interconnects and new improved instruction steering schemes. The results show that these point-to-point interconnects achieve much better performance than bus-based ones, and that the connectivity of the network together with effective steering schemes are key for high performance. We also show that these interconnects can be built with simple hardware and achieve a performance close to that of an idealized contention-free model. Joan-Manuel Parcerisa, Julio Sahuquillo, Antonio González 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | Characterizing the Dynamic Behavior of Workload Execution in SVM systemsabstractThe overhead associated with software management of shared virtual memory (SVM) systems can seriously impact overall system performance. One way to remedy this situation is to design more efficient SVM consistency protocols. In this paper we study a number of parallel workload characteristics that can negatively impact the performance of SVM systems. We attempt to quantify the sources of performance loss in some parallel workloads. Our goal is to better understand these characteristics, enabling us to develop SVM protocols that can adjust to dynamics in workload behavior. This paper has three main contributions: i) we measure the contention for synchronization resources, showing how applications exhibit distinct phases during their execution, ii) we quantify the relationship between page size and fragmentation/false sharing while varying the sharing unit size, and iii) we study the synergies between the contention for synchronization resources and fragmentation/false sharing, providing hints for developing improved protocols. Salvador Petit, Julio Sahuquillo, Ana Pont, David R. Kaeli |
SBAC-PAD | 2 |
| 2001 | About the sensitivity of the HLRC-DU protocol on diff and page sizesabstractRecent research on software distributed shared memory systems has focused on consistency protocols for improving performance. Home Lazy Release Consistency (HLRC) protocols have been widely adopted due to their performance advantages. Usually, these protocols invalidate pages through write notices. Variants of these protocols propose some criterion to update data of the corresponding pages instead of invalidating. In a previous paper, we proposed the HLRC-DU protocol, which is an improved version of the HLRC protocol. The HLRC-DU embeds update information in those write notices whose corresponding diff size is less than a given threshold, invalidating the remainder. The threshold trades off network bandwidth with update perfonnance. In this paper, we study the HLRC-DUprotocol’s sensitivity to page size and the threshold size selection. Our results show that while the page size slightly impacts performance, that our protocols are highly sensitive to the threshold value. Salvador Petit, Julio Sahuquillo, Ana Pont |
ISPASS | 2 |