VLDB 2026 Research / reviewers in the wild / expert
Salvador Petit
dblp:79/1638 · also Salvador Petit Marti
· DBLP profile ↗
89ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-2426-4134ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 76 · 6 first-author · 15 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SYNPA: understanding the impact of the interference modeling on thread-to-core allocation policies for SMT ARM processorsabstractModern high-performance servers increasingly rely on Simultaneous Multithreading (SMT) processors to enhance throughput with minimal area overhead. However, SMT architectures introduce inter-application interference, often resulting in degraded performance for individual applications. To address this issue, interference-aware thread-to-core (T2C) allocation policies are essential. This paper explores the design and implementation of such policies using real performance counters on ARM processors. We introduce the Instructions and Stalls Cycles (ISC) stack—a simple yet effective model for characterizing application behavior and identifying synergistic thread pairings. Building on our previous work, SYNPA, we improve the accuracy of the model by accounting for horizontal waste (that is, unused dispatch slots) and proposing methods to address limitations in ARM’s Performance Monitoring Unit (PMU), which prevent complete attribution of processor cycles. These enhancements result in a family of SYNPA schedulers, each based on a different ISC stack variant. Detailed discussions are provided on the pros and cons that researchers typically face when building a performance stack on commercial processors. These analyses are intended to assist researchers in their work. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
J. Supercomput. | 3 |
| 2026 | WAPA: A Microarchitecture- and Workload-Agnostic Universal SMT Scheduler
Marta Navarro 0001, Vicent Pallardó-Julià, Lucia Pons, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | WAPA: A Workload-Agnostic CPI-Based Thread-to-Core Allocation Policy
Marta Navarro 0001, Vicent Pallardó-Julià, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
Euro-Par (1) | 3 |
| 2025 | Power, energy, and performance analysis of single- and multi-threaded applications in the ARM ThunderX2abstractEnergy efficiency has been a major concern in data centers, and the problem is exacerbated as its size continues to rise. However, the lack of tools to measure and handle this energy at a fine granularity (e.g., processor core or last-level cache) has translated into slow research advances in this topic. Understanding where (i.e., which components) and when (the point in time) energy consumption translates into minor performance improvements is of paramount importance to design any energy-aware scheduler. This paper characterizes the relationship between energy consumption and performance in a 28-core ARM ThunderX2 processor for both single-threaded and multi-threaded applications. This paper shows that single-threaded applications with high CPU activity maintain their performance in spite of the inter-application interference at shared resources, but this comes at the expense of higher power consumption. Conversely, applications that heavily utilize the L3 cache and memory consume less power but suffer significant performance degradation as interference levels rise. In contrast, multi-threaded applications show two distinct behaviors. On the one hand, some of them experience significant performance gains when they execute in a higher number of cores with more threads, which outweighs the increase in power consumption, leading to high energy efficiency. Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
J. Parallel Distributed Comput. | 2 |
| 2025 | Advanced resource management: A hands-on master course in HPC and cloud computingabstractResource management has become a major concern in dealing with performance and fairness in recent computing servers, including a wide variety of shared resources. To achieve high-performing and efficient systems, both hardware and software engineers must be thoroughly trained in effective resource management techniques. This paper introduces the GRE master course (Spanish acronym for Resource Management and Performance Evaluation in Cloud and High-Performance Workloads), which is being offered since Fall 2023. The course is taught by instructors with broad research expertise in resource management and performance evaluation. Subjects covered in this course include workload characterization, state-of-the-art resource management approaches, and performance evaluation tools and methodologies used in production systems. Management techniques are studied both in the context of HPC and cloud computing, where resource efficiency is becoming a primary concern. To enhance the learning experience, the course integrates theoretical concepts with a wide set of hands-on tasks carried out on recent real platforms. A real cloud virtualized environment is mimicked using typical software deployed in production systems such as Proxmox Virtual Environment. Students learn to use tools such as Linux Perf and Intel Vtune Profiler, which are commonly employed by researchers and practitioners to carry out typical tasks like performance bottleneck analysis from a microarchitectural perspective. Overall, the GRE course provides students with a solid foundation and skills in resource management by addressing current hot topics both in the industry and academia. Student satisfaction and learning outcomes prove the success of the GRE course and encourage us to continue in this direction. Lucia Pons, Salvador Petit, Julio Sahuquillo |
J. Parallel Distributed Comput. | 2 |
| 2025 | Dual Fast-Track Cache: Organizing Ring-Shaped Racetracks to Work as L1 CachesabstractStatic Random-Access Memory (SRAM) is the fastest memory technology and has been the common design choice for implementing first-level (L1) caches in the processor pipeline, where speed is a key design issue that must be fulfilled. On the contrary, this technology offers much lower density compared to other technologies like Dynamic RAM, limiting L1 cache sizes of modern processors to a few tens of KB.This paper explores the use of slower but denser Domain Wall Memory (DWM) technology for L1 caches. This technology provides slow access times since it arranges multiple bits sequentially in a magnetic racetrack. To access these bits, they need to be shifted in order to place them under a header. A 1-bit shift usually takes one processor cycle, which can significantly hurt the application performance, making this working behavior inappropriate for L1 caches.Based on the locality (temporal and spatial) principles exploited by caches, this work proposes the Dual Fast-Track Cache (Dual FTC) design, a new approach to organizing a set of racetracks to build set-associative caches. Compared to a conventional SRAM cache, Dual FTC enhances storage capacity by a factor of 5 while incurring minimal shifting overhead, thereby rendering it a practical and appealing solution for L1 cache implementations.Experimental results show that the devised cache organization is as fast as an SRAM cache for 78% and 86% of the L1 data cache hits and L1 instruction cache hits, respectively (i.e., no shift is required). Consequently, due to the larger L1 cache capacities, significant system performance gains (by 22% on average) are obtained under the same silicon area. Alejandro Valero, Vicente Lorente, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 3 |
| 2024 | SYNPA: SMT Performance Analysis and Allocation of Threads to Cores in ARM ProcessorsabstractSimultaneous multithreading processors improve throughput over single-threaded processors thanks to sharing internal core resources among instructions from distinct threads. However, resource sharing introduces inter-thread interference within the core, which has a negative impact on individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the interference effects depends on the competing co-runners sharing the core. Thus, it can be mitigated by applying a thread-to-core allocation policy that smartly selects applications to be run in the same core to minimize their interference.This paper presents SYNPA, a simple approach that dynamically allocates threads to cores in an SMT processor based on their run-time dynamic behavior. The approach uses a regression model to select synergistic pairs to mitigate intra-core interference. The main novelty of SYNPA is that it uses just three variables collected from the performance counters available in current ARM processors at the dispatch stage. Experimental results show that SYNPA outperforms the default Linux scheduler by around 36%, on average, in terms of turnaround time in 8-application workloads combining frontend-bound and backend-bound benchmarks. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IPDPS | 3 |
| 2024 | Characterizing Power and Performance Interference Scalability in the 28-core ARM ThunderX2abstractNowadays, energy efficiency is a major concern in any type of processor-based device, ranging from processor servers to supercomputers, including mobile battery-fed devices. In recent years, systems based on the ARM architecture, tra-ditionally better suited for mobile and embedded systems, have increased their market share in segments commonly occupied by x86 processors. This growth is due, to some extent, to the excellent energy efficiency shown by high-performance ARM processors. Designing software and hardware energy-efficient systems requires a sound knowledge of the relationship among three main axes: component activity, power, and inter-application interference at the shared resources. This problem aggravates in many-core processors, which are ubiquitous in high-performance servers. This paper characterizes the aforementioned axes in a 28-core ARM Thunder X2 processor. Experimental results show that the performance of highly-scalable single-threaded applications is sustained regardless of the number of applications and interference introduced at the shared resources at the cost of increasing power. In contrast, low-scalable applications require much less power and experience a huge performance degradation as their number increases. Ibai Calero, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PDP | 2 |
| 2024 | A modular approach to build a hardware testbed for cloud resource management research
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo |
J. Supercomput. | 2 |
| 2023 | Thread-to-Core Allocation in ARM Processors Building Synergistic PairsabstractSimultaneous multithreading (SMT) processors can present significant throughput improvements over single-threaded (ST) processors thanks to sharing internal core resources among instructions executing from multiple threads. However, resource sharing introduces inter-thread interference within the core, which negatively impacts individual application performance and can significantly increase the turnaround time of multi-program workloads. The severity of the intra-core interference on performance depends on the applications co-running in the same core. A thread scheduler can help reduce this effect by smartly selecting the pairs of applications that should run on each SMT core. This paper presents SYNPA, a simple approach that dynamically allocates threads to SMT cores based on their run-time dynamic behavior. SYNPA uses a regression model to select synergistic pairs to mitigate intra-core interference. Results show that SYNPA outperforms the default Linux scheduler by around 35%, on average, in terms of turnaround time when running 8-application workloads combining frontend-bound and backend-bound applications. Marta Navarro 0001, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PACT | 3 |
| 2023 | Stratus: A Hardware/Software Infrastructure for Controlled Cloud ResearchabstractCloud systems deploy a wide variety of shared resources and host a large number of tenant applications. To perform cloud research, a small experimental platform is commonly used, which hides the huge system complexity and provides flexibility. Despite being simpler, this platform should include the main cloud system components (hardware and software) to provide representative results. A wide set of platforms have spread in recent years; however, most of them only include a major cloud component or lack the deployment of virtual machines (VMs) to provide isolation. This paper presents Stratus, an experimental platform that is currently being used to carry out cloud research. To the best of our knowledge, Stratus is the only platform that jointly provides three main features: uses VMs to isolate tenant applications, deploys the three types of cloud nodes (server, client, and storage), and manages all main shared system resources (CPUs, LLC space, memory, network, and disk bandwidth). Moreover, Stratus implements a software manager to ease the research and aid the design of QoS-aware policies. The manager integrates three main functionalities: management and control of the execution of VMs and running applications, monitoring of hardware performance counters and system resource utilization, and partitioning of the main shared system resources by using technologies available in commercial processors. Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Chaoyi Huang, Julio Sahuquillo |
PDP | 2 |
| 2023 | Cloud White: Detecting and Estimating QoS Degradation of Latency-Critical Workloads in the Public CloudabstractThe increasing popularity of cloud computing has forced cloud providers to build economies of scale to meet the growing demand. Nowadays, data-centers include thousands of physical machines, each hosting many virtual machines (VMs), which share the main system resources, causing interference that can significantly impact on performance. Frequently, these data-centers run latency-critical workloads, whose performance is determined by tail latency, which is very sensitive to the interference of co-running workloads. To prevent QoS violations, cloud providers adopt overprovisioning strategies but they reduce the server utilization and increase the costs. A mechanism that accurately estimates performance degradation dynamically in a production system would allow cloud providers to improve the servers’ utilization. In this work we propose Cloud White, an approach that is able to detect the inter-VM interference in scenarios with multiple co-located latency-critical VMs and estimate the performance degradation using multi-variable regression models. Unlike previous proposals, Cloud White is built taking into account the limitations of a public cloud production system. Experimental results show that Cloud White is able to estimate performance degradation with a small overall prediction error of 5%. Lucia Pons, Josué Feliu, Julio Sahuquillo, María Engracia Gómez, Salvador Petit, Julio Pons, Chaoyi Huang |
Future Gener. Comput. Syst. | 5 |
| 2022 | Cache-Poll: Containing Pollution in Non-Inclusive Caches Through Cache PartitioningabstractCurrent server processors have redistributed the cache hierarchy space over previous generations. The private L2 cache has been made larger and the shared last level caches (LLC) smaller but designed as non-inclusive to reduce the number of replicated blocks. As a result, the new organization shrinks the per-core cache area. Lucia Pons, Julio Sahuquillo, Salvador Petit, Julio Pons |
ICPP | 3 |
| 2022 | Fast-track cache: a huge racetrack memory L1 data cacheabstractFirst-level (L1) caches have been traditionally implemented with Static Random-Access Memory (SRAM) technology, since it is the fastest memory technology, and L1 caches call for tight timing constraints in the processor pipeline. However, one of the main downsides of SRAM is its low density, which prevents L1 caches to improve their storage capacity beyond a few tens of KB. On the other hand, the recent Domain Wall Memory (DWM) technology overcomes such a constraint by arranging multiple bits in a magnetic racetrack, and sharing a header to access those bits. Accessing a bit requires a shift operation to align the target bit under the header. Such shifts increase the final access latency, which is the main reason why DWM has been mostly used to implement slow last-level caches. Hugo Tárrega, Alejandro Valero, Vicente Lorente, Salvador Petit, Julio Sahuquillo |
ICS | 4 |
| 2022 | A Neural Network to Estimate Isolated Performance from Multi-Program ExecutionabstractWhen multiple applications are running on a platform with shared resources like multicore CPUs, the behaviour of the running application can be altered by the co-runners. In this case, the system resources need to be managed (e.g. by repartitioning the cache space, re-schedule applications in distinct cores, modifying the prefetcher configuration, etc.) to reduce the inter-application interference in order to minimize the performance losses over isolated execution. In this context, a main challenge in different computing scenarios like the public cloud or soft real-time systems is knowing the performance impact of a given management action on each application with respect to its isolated execution. With this aim, in this work we present a neural network-based approach that estimates the performance an application would have had in isolation from multi-program executions. Experimental results show that the proposal dynamically adapts to changes in application behavior. On average, the predicted performance presents an error deviation by 11.7% and 2.3% for MAPE and MSE respectively. Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
PDP | 3 |
| 2022 | Effect of Hyper-Threading in Latency-Critical Multithreaded Cloud Applications and Utilization Analysis of the Major System ResourcesabstractMultithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main system resources is a major concern for cloud providers in order to devise virtualization strategies that improve the system efficiency. With this aim, this paper first characterizes the impact of QPS on tail latency, analyzing different scenarios varying the number of threads and the thread-to-core allocation (single-task and multi-task execution) policy. The characterization study reveals that the performance of some applications does not scale with the number of threads, and the performance of some others is insensitive to the Hyper-Threading technology, so they can be allocated in less physical cores and improve system utilization. Identifying these applications, however, at run-time is challenging. Despite identifying these applications at run-time is challenging, this paper shows that they can be successfully detected at run-time by analyzing the utilization trend of the major system resources. In addition to CPU, we have also studied how assigning the share of each application of other major shared system resources impacts on performance. We outline considerations cloud providers should take into account to improve performance and resource utilization. Lucia Pons, Josué Feliu, José Puche, Chaoyi Huang, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo |
Future Gener. Comput. Syst. | 5 |
| 2022 | VMT: Virtualized Multi-Threading for Accelerating Graph Workloads on Commodity ProcessorsabstractModern-day graph workloads operate on huge graphs through pointer chasing which leads to high last-level cache (LLC) miss rates and limited memory-level parallelism (MLP). Simultaneous Multi-Threading (SMT) effectively hides the memory access latencies for multi-threaded graph workloads provided that sufficient threads are supported in hardware. Unfortunately, providing a sufficiently large number of physical threads incurs an unjustifiably high hardware cost for commodity SMT processors which typically implement only two physical hardware threads. Ideally, we would like to achieve aggressive-SMT performance when running graph workloads on modest commodity processors. In this paper, we propose Virtualized Multi-Threading (VMT), a low-overhead multi-threading paradigm for accelerating graph workloads on commodity processors. Unlike prior multi-threading paradigms, VMT virtualizes both the physical hardware threads and the architecture state: VMT maps a large number of logical software threads to a small number of physical hardware threads, while maintaining the architecture state of the logical threads in the processor's cache hierarchy. Implemented on top of a quad-core 2-way SMT processor, VMT achieves an average speedup of 1.74× for a set of representative graph workloads, while incurring minimal hardware cost (195 bytes per core to support up to 32 logical threads). VMT's low hardware cost paves the way for implementation in commodity processors. Josué Feliu, Ajeya Naithani, Julio Sahuquillo, Salvador Petit, Moinuddin K. Qureshi, Lieven Eeckhout |
IEEE Trans. Computers | 4 |
| 2022 | DeepP: Deep Learning Multi-Program Prefetch Configuration for the IBM POWER 8abstractCurrent multi-core processors implement sophisticated hardware prefetchers, that can be configured by application (PID), to improve the system performance. When running multiple applications, each application can present different prefetch requirements, hence different configurations can be used. Setting the optimal prefetch configuration for each application is a complex task since it does not only depend on the application characteristics but also on the interference at the shared memory resources (e.g., memory bandwidth). In his paper, we proposeDeepP, a deep learning approach for the IBM POWER8 that identifies at run-time the best prefetch configuration for each application in a workload. To this end, the neural network predicts the performance of each application under the studied prefetch configurations by using a set of performance events. The prediction accuracy of the network is improved thanks to a dynamic training methodology that allows learning the impact of dynamic changes of the prefetch configuration on performance. At run-time, the devised network infers the best prefetch configuration for each application and adjusts it dynamically. Experimental results show that the proposed approach improves performance, on average, by 5.8%, 6.7%, and 15.8% compared to the default prefetch configuration across different 6-, 8-, and 10-application workloads, respectively. Manel Lurbe, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Computers | 3 |
| 2020 | Impact of the Array Shape and Memory Bandwidth on the Execution Time of CNN Systolic ArraysabstractThe use of Convolutional Neural Networks (CNN) has experienced a huge rise over the last recent years and its popularity has increased exponentially, mainly due to its application both for image recognition and certain applications related to artificial intelligence. The new applications of CNN request computing demands that are difficult to address by conventional processors.As a consequence, accelerators -both prototypes and commercial products- focusing on CNN computation have been proposed. Among these accelerators, those based on systolic arrays have acquired a special relevance; some examples are the Google's TPU and Eyeriss.Current research has focused on regular squared systolic arrays and most existing work assumes that there is enough memory bandwidth to feed the systolic array with input data. In this paper we explore the design of non-squared systolic arrays and address the impact of the memory bandwidth from a performance perspective.This work makes two main contributions. First, we found that some workloads with non-squared arrays achieve similar performance to systolic arrays twice as large, which can translate in area and/or energy benefits.Second, we present a performance comparison varying the main memory bandwidth for current DRAM devices. The analysis reveals that main memory bandwidth has a great impact on performance and that the decision of which technology use is key for the system performance. For the 64x64 array size it is necessary to use HBM2 memory to avoid the slowdown that would introduce cheaper technologies (e.g. DDR5 and DDR4). Eduardo Yago, Pau Castelló, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
DSD | 3 |
| 2020 | An efficient cache flat storage organization for multithreaded workloads for low power processors
José Puche, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
Future Gener. Comput. Syst. | 2 |
| 2020 | Thread Isolation to Improve Symbiotic Scheduling on SMT Multicore ProcessorsabstractResource sharing is a critical issue in simultaneous multithreading (SMT) processors as threads running simultaneously on an SMT core compete for shared resources. Symbiotic job scheduling, which co-schedules applications with complementary resource demands, is an effective solution to maximize hardware utilization and improve overall system performance. However, symbiotic job scheduling typically distributes threads evenly among cores, i.e., all cores get assigned the same number of threads, which we find to lead to sub-optimal performance. In this paper, we show that asymmetric schedules (i.e., schedules that assign a different number of threads to each SMT core) can significantly improve performance compared to symmetric schedules. To leverage this finding, we propose thread isolation, a technique that turns symmetric schedules into asymmetric ones yielding higher overall system performance. Thread isolation identifies SMT-adverse applications and schedules them in isolation on a dedicated core to mitigate their sharp performance degradation under SMT. Our experimental results on an IBM POWER8 processor show that thread isolation improves system throughput by up to 5.5 percent compared to a state-of-the-art symmetric symbiotic job scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, Lieven Eeckhout |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Bandwidth-Aware Dynamic Prefetch Configuration for IBM POWER8abstractAdvanced hardware prefetch engines are being integrated in current high-performance processors. Prefetching can boost the performance of most applications, however, the induced bandwidth consumption can lead the system to a high contention for main memory bandwidth, which is a scarce resource in current multicores. In such a case, the system performance can be severely damaged. This article characterizes the applications’ behavior in an IBM POWER8 machine, which presents many prefetch settings, varying the bandwidth contention. The study reveals that the best prefetch setting for each application depends on the main memory bandwidth availability, that is, it depends on the co-running applications. Based on this study, we propose Bandwidth-Aware Prefetch Configuration (BAPC) a scalable adaptive prefetching algorithm that improves the performance of multi-program workloads. BAPC increases the performance of the applications in a 12, 15, and 16 percent of 6-, 8-, and 10-application workloads over the IBM POWER8 default configuration. In addition, BAPC reduces bandwidth consumption in 39, 42, and 45 percent, respectively. Carlos Navarro, Josué Feliu, Salvador Petit, María Engracia Gómez, Julio Sahuquillo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Phase-Aware Cache Partitioning to Target Both Turnaround Time and System PerformanceabstractThe Last Level Cache (LLC) plays a key role in the system performance of current multi-cores by reducing the number of long latency main memory accesses. The inter-application interference at this shared resource, however, can lead the system to undesired situations regarding performance and fairness. Recent approaches have successfully addressed fairness and turnaround time (TT) in commercial processors. Nevertheless, these approaches must face sustaining system performance, which is challenging. This work makes two main contributions. LLC behaviors regarding cache performance, data reuse and cache occupancy, that adversely impact on the final performance are identified. Second, based on these behaviors, we propose the Critical-Phase Aware Partitioning Approach (CPA), which reduces TT while sustaining (and even improving) IPC by making an effective use of the LLC space. Experimental results show that CPA outperforms CA, Dunn and KPart state-of-the-art approaches, and improves TT (over 40 percent in some workloads) over Linux default behavior while sustaining or even improving IPC by more than 3 percent in several mixes. Lucia Pons, Julio Sahuquillo, Vicent Selfa, Salvador Petit, Julio Pons |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Modeling and analysis of the performance of exascale photonic networksabstractSummary Photonics technology has become a promising and viable alternative for both on‐chip and off‐chip interconnection networks of future Exascale systems. Nevertheless, this technology is not mature enough yet in this context, so research efforts focusing on photonic networks are still required to achieve realistic suitable network implementations. In this regard, system‐level photonic network simulators can help guide designers to assess the multiple design choices. Most current research is done on electrical network simulators, whose components work widely different from photonics components. In this work, we summarize and compare the working behavior of both technologies which includes the use of optical routers, wavelength‐division multiplexing and circuit switching among others. After implementing them into a well‐known simulation framework, an extensive simulation study has been carried out using realistic photonic network configurations with synthetic and realistic traffic. Experimental results show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. Our study also reveals that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Jose Duro, Jose Antonio Pascual, Salvador Petit, Julio Sahuquillo, María Engracia Gómez |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Efficient Management of Cache Accesses to Boost GPGPU Memory Subsystem PerformanceabstractTo support the massive amount of memory accesses that GPGPU applications generate, GPU memory hierarchies are becoming more and more complex, and the Last Level Cache (LLC) size considerably increases each GPU generation. This paper shows that counter-intuitively, enlarging the LLC brings marginal performance gains in most applications. In other words, increasing the LLC size does not scale neither in performance nor energy consumption. We examine how LLC misses are managed in typical GPUs, and we find that in most cases the way LLC misses are managed are precisely the main performance limiter. This paper proposes a novel approach that addresses this shortcoming by leveraging a tiny additional Fetch and Replacement Cache-like structure (FRC) that stores control and coherence information of the incoming blocks until they are fetched from main memory. Then, the fetched blocks are swapped with the victim blocks (i.e., selected to be replaced) in the LLC, and the eviction of such victim blocks is performed from the FRC. This approach improves performance due to three main reasons: i) the lifetime of blocks being replaced is enlarged, ii) the main memory path is unclogged on long bursts of LLC misses, and iii) the average LLC miss latency is reduced. The proposal improves the LLC hit ratio, memory-level parallelism, and reduces the miss latency compared to much larger conventional caches. Moreover, this is achieved with reduced energy consumption and with much less area requirements. Experimental results show that the proposed FRC cache scales in performance with the number of GPU compute units and the LLC size, since, depending on the FRC size, performance improves ranging from 30 to 67 percent for a modern baseline GPU card, and from 32 to 118 percent for a larger GPU. In addition, energy consumption is reduced on average from 49 to 57 percent for the larger GPU. These benefits come with a small area increase (by 7.3 percent) over the LLC baseline. Francisco Candel, Alejandro Valero, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 3 |
| 2019 | An Aging-Aware GPU Register File Design Based on Data RedundancyabstractNowadays, GPUs sit at the forefront of high-performance computing thanks to their massive computational capabilities. Internally, thousands of functional units, architected to be fed by large register files, fuel such a performance. At deep nanometer technologies, the SRAM memory cells that implement GPU register files are very sensitive to the Negative Bias Temperature Instability (NBTI) effect. NBTI ages cell transistors by degrading their threshold voltage$V_{th}$over the lifetime of the GPU. This degradation, which manifests when a cell keeps the same logic value for a relatively long period of time, compromises the cell read stability and increases the transistor switching delay, which can lead to wrong read values and eventually exceed the processor cycle time, respectively, so resulting in faulty operation. This work proposes architectural mechanisms leveraging the redundancy of the data stored in GPU register files to attack NBTI aging. The proposed mechanisms are based on data compression, power gating, and register address rotation techniques. All these mechanisms working together balance the distribution of logic values stored in the cells along the execution time, reducing both the overall$V_{th}$degradation and the increase in the transistor switching delays. Experimental results show that a conventional GPU register file suffers the worst case for NBTI, since a significant fraction of the cells maintain the same logic value during the entire application execution (i.e., a 100 percent ‘0’ and ‘1’ duty cycle distributions). On average, the proposal reduces these distributions by 58 and 68 percent, respectively, which translates into$V_{th}$degradation savings by 54 and 62 percent, respectively. Alejandro Valero, Francisco Candel, Darío Suárez Gracia, Salvador Petit, Julio Sahuquillo |
IEEE Trans. Computers | 4 |
| 2019 | FOS: a low-power cache organization for multicores
José Puche, Salvador Petit, Julio Sahuquillo, María Engracia Gómez |
J. Supercomput. | 2 |
| 2019 | Way Combination for an Adaptive and Scalable Coherence DirectoryabstractThis manuscript opens the way to a new class of coherence directory structures that are based on the brand-new concept of way combining. A Way-Combining Directory (WC-dir) builds on a typical sparse directory but allows to take advantage of several ways in the same set to codify the sharing information of each memory block. The result is a sparse directory with variable effective associativity per set and variable length entries, thus being able to dynamically adapt the directory structure to the particular requirements of each application. In particular, our proposal uses just enough bits per entry to store a single pointer, which is optimal for the common case of having just one sharer. For those addresses that have more than one sharer, we have observed that in the majority of cases extra bits could be taken from other empty ways in the same set. All in all, our proposal minimizes the storage overheads without losing the flexibility to adapt to several sharing degrees and without the complexities of other previously proposed techniques. Detailed simulations of a 128-core multicore architecture running benchmarks from PARSEC-3.0 and SPLASH-3 demonstrate that WC-dir can closely approach the performance of a non-scalable bit vector sparse directory, beating the state-of-the-art Scalable Coherence Directory (SCD) and Pool directory proposals. J. Rubén Titos Gil, Antonio Flores, Ricardo Fernández-Pascual, Alberto Ros 0001, Salvador Petit, Julio Sahuquillo, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Improving GPU Cache Hierarchy Performance with a Fetch and Replacement Cache
Francisco Candel, Salvador Petit, Alejandro Valero, Julio Sahuquillo |
Euro-Par | 2 |
| 2018 | Improving System Turnaround Time with Intel CAT by Identifying LLC Critical Applications
Lucia Pons, Vicent Selfa, Julio Sahuquillo, Salvador Petit, Julio Pons |
Euro-Par | 4 |
| 2018 | Accurately modeling the on-chip and off-chip GPU memory subsystem
Francisco Candel, Salvador Petit, Julio Sahuquillo, José Duato |
Future Gener. Comput. Syst. | 2 |
| 2018 | Designing lab sessions focusing on real processors for computer architecture courses: A practical perspective
Josué Feliu, Julio Sahuquillo, Salvador Petit |
J. Parallel Distributed Comput. | 3 |
| 2017 | Application Clustering Policies to Address System Fairness with Intel's Cache Allocation TechnologyabstractAchieving system fairness is a major design concern in current multicore processors. Unfairness arises due to contention in the shared resources of the system, such as the LLC and main memory. To address this problem, many research works have proposed novel cache partitioning policies aimed at addressing system fairness without harming performance. Unfortunately, existing proposals targeting fairness require extra hardware which makes them impractical in commercial processors.Recent Intel Xeon processors feature Cache Allocation Technology (CAT), a hardware cache partitioning mechanism that can be controlled from userspace software and that allows to create partitions in the LLC and assign different groups of applications to them.In this paper we propose a family of clustering-based cache partitioning policies to address fairness in systems that feature Intel's CAT. The proposal acts at two levels: applications showing similar amount of core stalls due to LLC accesses are first grouped into clusters, after which each cluster is given a number of ways using a simple mathematical model. To the best of our knowledge, this is the first attempt to address system fairness using the cache partitioning hardware in a real product. Results show that our best performing policy reduces system unfairness by up to 80% (39% on average) for 8-application workloads and by up to 45% (25% on average) for 12-application workloads compared to a non-partitioning approach. Vicent Selfa, Julio Sahuquillo, Lieven Eeckhout, Salvador Petit, María Engracia Gómez |
PACT | 4 |
| 2017 | Exploiting Data Compression to Mitigate Aging in GPU Register FilesabstractNowadays, GPUs sit at the forefront of highperformance computing thanks to their massive computational capabilities. Internally, thousands of functional units, architected to be fed by large register files, fuel such a performance.At nanometer technologies, the SRAM cells that implement register files suffer the Negative Bias Temperature Instability (NBTI) effect, which degrades the transistor threshold voltage Vth and, in turn, can make cells faulty unreliable when they hold the same logic value for long periods of time.Fortunately, the GPU single-thread multiple-data execution model writes data in recognizable patterns. This work proposes mechanisms to detect those patterns, and to compress and shuffle the data, so that compressed register file entries can be safely powered off, mitigating NBTI aging.Experimental results show that a conventional GPU register file experiences the worst case for NBTI, since maintains cells with a single logic value during the entire application execution (i.e., a 100% 0 and 1 duty cycle distributions). On average, the proposal reduces these distributions by 61% and 72%, respectively, which translates into Vth degradation savings by 57% and 64%, respectively. Francisco Candel, Alejandro Valero, Salvador Petit, Darío Suárez Gracia, Julio Sahuquillo |
SBAC-PAD | 3 |
| 2017 | A research-oriented course on Advanced Multicore Architecture: Contents and active learning methodologies
Salvador Petit, Julio Sahuquillo, María Engracia Gómez, Vicent Selfa |
J. Parallel Distributed Comput. | 1 |
| 2017 | Perf&Fair: A Progress-Aware Scheduler to Enhance Performance and Fairness in SMT MulticoresabstractNowadays, high performance multicore processors implement multithreading capabilities. The processes running concurrently on these processors are continuously competing for the shared resources, not only among cores, but also within the core. While resource sharing increases the resource utilization, the interference among processes accessing the shared resources can strongly affect the performance of individual processes and its predictability. In this scenario, process scheduling plays a key role to deal with performance and fairness. In this work we present a process scheduler for SMT multicores that simultaneously addresses both performance and fairness. This is a major design issue since scheduling for only one of the two targets tends to damage the other. To address performance, the scheduler tackles bandwidth contention at the L1 cache and main memory. To deal with fairness, the scheduler estimates the progress experienced by the processes, and gives priority to the processes with lower accumulated progress. Experimental results on an Intel Xeon E5645 featuring six dual-threaded SMT cores show that the proposed scheduler improves both performance and fairness over two state-of-the-art schedulers and the Linux OS scheduler. Compared to Linux, unfairness is reduced to a half while still improving performance by 5.6 percent. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IEEE Trans. Computers | 3 |
| 2017 | Improving IBM POWER8 Performance Through Symbiotic Job SchedulingabstractSymbiotic job scheduling, i.e., scheduling applications that co-run well together on a core, can have a considerable impact on the performance of processors with simultaneous multithreading (SMT) cores. SMTcores share most of their microarchitectural components among the co-running applications, which causes performance interference between them. Therefore, scheduling applications with complementary resource requirements on the same core can greatly improve the throughput of the system. This paper enhances symbiotic job scheduling for the IBM POWER8 processor. We leverage the existing cycle accounting mechanism to build an interference model that predicts symbiosis between applications. The proposed models achieve higher accuracy than previous models by predicting job symbiosis from throttled CPI stacks, i.e., CPI stacks of the applications when running in the same SMT mode to consider the statically partitioned resources, but without interference from other applications. The symbiotic scheduler uses these interference models to decide, at run-time, which applications should run on the same core or on separate cores. We prototype the symbiotic scheduler as a user-level scheduler in the Linux operating system and evaluate it on an IBM POWER8 server running multiprogram workloads. The symbiotic job scheduler significantly improves performance compared to both an agnostic random scheduler and the default Linux scheduler. Across all evaluated workloads in SMT4 mode, throughput improves by 12.4 and 5.1 percent on average over the random and Linux schedulers, respectively. Josué Feliu, Stijn Eyerman, Julio Sahuquillo, Salvador Petit, Lieven Eeckhout |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | A Hardware Approach to Fairly Balance the Inter-Thread Interference in Shared CachesabstractShared caches have become the common design choice in the vast majority of modern multi-core and many-core processors, since cache sharing improves throughput for a given silicon area. Sharing the cache, however, has a downside: the requests from multiple applications compete among them for cache resources, so the execution time of each application increases over isolated execution. The degree in which the performance of each application is affected by the interference becomes unpredictable yielding the system to unfairness situations. This paper proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead hardware-based cache partitioning approach that addresses system fairness. FPCP reduces the interference by allocating to each application a cache partition and adjusting the partition sizes at runtime. To adjust partitions, our approach estimates during multicore execution the time each application would have taken in isolation, which is challenging. The proposed approach has two main differences over existing approaches. First, FPCP distributes cache ways incrementally, which makes the proposal less prone to estimation errors. Second, the proposed algorithm is much less costly than the state-of-the-art ASM-Cache approach. Experimental results show that, compared to ASM-Cache, FPCP reduces unfairness by 48 percent in four-application workloads and by 28 percent in eight-application workloads, without harming the performance. Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | On Microarchitectural Mechanisms for Cache Wearout ReductionabstractHot carrier injection (HCI) and bias temperature instability (BTI) are two of the main deleterious effects that increase a transistor's threshold voltage over the lifetime of a microprocessor. This voltage degradation causes slower transistor switching and eventually can result in faulty operation. HCI manifests itself when transistors switch from logic “0” to “1” and vice versa, whereas BTI is the result of a transistor maintaining the same logic value for an extended period of time. These failure mechanisms are especially acute in those transistors used to implement the SRAM cells of first-level (L1) caches, which are frequently accessed, so they are critical to performance, and they are continuously aging. This paper focuses on microarchitectural solutions to reduce transistor aging effects induced by both HCI and BTI in the data array of L1 data caches. First, we show that the majority of cell flips are concentrated in a small number of specific bits within each data word. In addition, we also build upon the previous studies, showing that logic “0” is the most frequently written value in a cache by identifying which cells hold a given logic value for a significant amount of time. Based on these observations, this paper introduces a number of architectural techniques that spread the number of flips evenly across memory cells and reduce the amount of time that logic “0” values are stored in the cells by switching OFF specific data bytes. Experimental results show that the threshold voltage degradation savings range from 21.8% to 44.3% depending on the application. Alejandro Valero, Negar Miralaei, Salvador Petit, Julio Sahuquillo, Timothy M. Jones 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Student Research Poster: A Low Complexity Cache Sharing Mechanism to Address System FairnessabstractShared caches have become, de facto, the common design choice in current multi-cores, ranging from embedded devices to high-performance processors. In these systems, requests from multiple applications compete for the cache resources, degrading to different extents their progress, quantified as the performance of individual applications compared to isolated execution. The difference between the progresses of the running applications yields the system to unpredictable behavior and causes a fairness problem. This problem can be addressed by carefully partitioning cache resources among the contending applications, but to be effective, a partitioning approach needs to estimate per-application progress. This work proposes Fair-Progress Cache Partitioning (FPCP), a low-overhead cache partitioning approach which addresses fairness by distributing cache resources among applications depending on their progress. To estimate progress, we have implemented two state-of-the-art performance models, ASM and PTCA, which estimate, at runtime, the performance a given application would have if executed in isolation. Vicent Selfa, Julio Sahuquillo, Salvador Petit, María Engracia Gómez |
PACT | 3 |
| 2016 | Symbiotic job scheduling on the IBM POWER8abstractSimultaneous multithreading (SMT) processors share most of the microarchitectural core components among the co-running applications. The competition for shared resources causes performance interference between applications. Therefore, the performance benefits of SMT processors heavily depend on the complementarity of the co-running applications. Symbiotic job scheduling, i.e., scheduling applications that co-run well together on a core, can have a considerable impact on the performance of a processor with SMT cores. Prior work uses sampling or novel hardware support to perform symbiotic job scheduling, which has either a non-negligible overhead or is impossible to use on existing hardware. This paper proposes a symbiotic job scheduler for the IBM POWER8 processor. We leverage the existing cycle accounting mechanism to predict symbiosis between applications, and use that information at run-time to decide which applications should run on the same core or on separate cores. We implement the scheduler in the Linux operating system and evaluate it on an IBM POWER8 server running multiprogrammed workloads. The symbiotic job scheduler significantly improves performance compared to both an agnostic random scheduler and the default Linux scheduler. With respect to Linux, it achieves an average speedup by 8.8% for workloads comprising 12 applications, and by 4.7% on average across all evaluated workloads. Josué Feliu, Stijn Eyerman, Julio Sahuquillo, Salvador Petit |
HPCA | 4 |
| 2016 | Impact of Memory-Level Parallelism on the Performance of GPU Coherence ProtocolsabstractGraphics Processing Units (GPUs) are being implemented in heterogeneous CPU/GPU systems due their high efficiency when executing massively parallel applications. New challenges appear to deal with heterogenous coherence in these systems due to the huge amount (hundreds or thousands) of on-going memory requests of GPUs, which is limited by the Miss Status Holding Register (MSHR) file size associated to the L1 cache. This paper analyzes how the number of MSHRs i) affects to typical memory performance metrics and ii) impacts on the system performance under two recent GPU coherence protocols, called NMOESI and SI (Southern Islands), which introduce distinct coherence traffic. We find two key findings that can help improve the performance of coherence protocols. First, there is a strong correlation between system performance and memory subsystem latency regardless of the used protocol. Second, system performance varies with the number of supported cache misses, however, counterintuitively, supporting more cache misses does not always bring enhanced performance but it can turn into performance drops. Francisco Candel, Salvador Petit, Julio Sahuquillo, José Duato |
PDP | 2 |
| 2016 | A dynamic execution time estimation model to save energy in heterogeneous multicores running periodic tasks
Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Luis March, José Duato |
Future Gener. Comput. Syst. | 3 |
| 2016 | Bandwidth-Aware On-Line Scheduling in SMT MulticoresabstractThe memory hierarchy plays a critical role on the performance of current chip multiprocessors. Main memory is shared by all the running processes, which can cause important bandwidth contention. In addition, when the processor implements SMT cores, the L1 bandwidth becomes shared among the threads running on each core. In such a case, bandwidth-aware schedulers emerge as an interesting approach to mitigate the contention. This work investigates the performance degradation that the processes suffer due to memory bandwidth constraints. Experiments show that main memory and L1 bandwidth contention negatively impact the process performance; in both cases, performance degradation can grow up to 40 percent for some of applications. To deal with contention, we devise a scheduling algorithm that consists of two policies guided by the bandwidth consumption gathered at runtime. The process selection policy balances the number of memory requests over the execution time to address main memory bandwidth contention. The process allocation policy tackles L1 bandwidth contention by balancing the L1 accesses among the L1 caches. The proposal is evaluated on a Xeon E5645 platform using a wide set of multiprogrammed workloads, achieving performance benefits up to 6.7 percent with respect to the Linux scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IEEE Trans. Computers | 3 |
| 2015 | Addressing Fairness in SMT Multicores with a Progress-Aware SchedulerabstractCurrent SMT (simultaneous multithreading) processors co-schedule jobs on the same core, thus sharing core resources like L1 caches. In SMT multicores, threads also compete among themselves for uncore resources like the LLC (last level cache) and DRAM modules. Per process performance degradation over isolated execution mainly depends on process resource requirements and the resource contention induced by co-runners. Consequently, the running processes progress at different pace. If schedulers are not progress aware, the unpredictable execution time caused by unfairness can introduce undesirable behaviors on the system such as difficulties to keep priority-based scheduling. This work proposes a job scheduler for SMT multicores that provides fairness to the execution of multi programmed workloads. To this end, the scheduler estimates per-process standalone performance by periodically creating low-contention co-schedules. These estimates are used to compute the per process progress. Then, those processes with less progress are prioritized to enhance fairness. Experimental results on a Intel Xeon with six dual-threaded SMT cores show that the proposed scheduler reduces unfairness, on average, by 3× over Linux OS. Moreover, thanks to the tread to core allocation policy, the scheduler slightly improves throughput and turnaround time. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IPDPS | 3 |
| 2015 | Design of Hybrid Second-Level CachesabstractIn recent years, embedded dynamic random-access memory (eDRAM) technology has been implemented in last-level caches due to its low leakage energy consumption and high density. However, the fact that eDRAM presents slower access time than static RAM (SRAM) technology has prevented its inclusion in higher levels of the cache hierarchy. This paper proposes to mingle SRAM and eDRAM banks within the data array of second-level (L2) caches. The main goal is to achieve the best trade-off among performance, energy, and area. To this end, two main directions have been followed. First, this paper explores the optimal percentage of banks for each technology. Second, the cache controller is redesigned to deal with performance and energy. Performance is addressed by keeping the most likely accessed blocks in fast SRAM banks. In addition, energy savings are further enhanced by avoiding unnecessary destructive reads of eDRAM blocks. Experimental results show that, compared to a conventional SRAM L2 cache, a hybrid approach requiring similar or even lower area speedups the performance on average by 5.9 percent, while the total energy savings are by 32 percent. For a 45 nm technology node, the energy-delay-area product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless of the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is an eighth of the total number of cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 3 |
| 2014 | Addressing bandwidth contention in SMT multicores through schedulingabstractTo mitigate the impact of bandwidth contention, which in some processes can yield to performance degradations up to 40%, we devise a scheduling algorithm that tackles main memory and L1 bandwidth contention. Experimental evaluation on a real system shows that the proposal achieves an average speedup by 5% with respect to Linux. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
ICS | 3 |
| 2014 | Cache-Hierarchy Contention-Aware Scheduling in CMPsabstractTo improve chip multiprocessor (CMP) performance, recent research has focused on scheduling strategies to mitigate main memory bandwidth contention. Nowadays, commercial CMPs implement multilevel cache hierarchies that are shared by several multithreaded cores. In this microprocessor design, contention points may appear along the whole memory hierarchy. Moreover, this problem is expected to aggravate in future technologies, since the number of cores and hardware threads, and consequently the size of the shared caches increase with each microprocessor generation. This paper characterizes the impact on performance of the different contention points that appear along the memory subsystem. The analysis shows that some benchmarks are more sensitive to contention in higher levels of the memory hierarchy (e.g., shared L2) than to main memory contention. In this paper, we propose two generic scheduling strategies for CMPs. The first strategy takes into account the available bandwidth at each level of the cache hierarchy. The strategy selects the processes to be coscheduled and allocates them to cores to minimize contention effects. The second strategy also considers the performance degradation each process suffers due to contention-aware scheduling. Both proposals have been implemented and evaluated in a commercial single-threaded quad-core processor with a relatively small two-level cache hierarchy. The proposals reach, on average, a performance improvement by 5.38 and 6.64 percent when compared with the Linux scheduler, while this improvement is by 3.61 percent for an state-of-the-art memory contention-aware scheduler under the evaluated mixes. Josué Feliu, Salvador Petit, Julio Sahuquillo, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Efficient Register Renaming and Recovery for High-Performance ProcessorsabstractModern superscalar processors implement register renaming using either random access memory (RAM) or content-addressable memories (CAM) tables. The design of these structures should address both access time and misprediction recovery penalty. Although direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. The presence of associative ports in CAMs, however, prevents them from scaling with the number of physical registers and pipeline width, negatively impacting performance, area, and energy consumption at the rename stage. In this paper, we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides fast and energy-efficient access to register mappings. On misspeculation, a low-complexity CAM enables immediate recovery. Experimental results show that in a four-way state-of-the-art superscalar processor, the new approach provides almost the same performance as an ideal CAM-based renaming scheme, while dissipating only between 17% and 26% of the original energy and, in some cases, consuming less energy than purely RAM-based renaming schemes. Overall, the silicon area required to implement the hybrid RAM-CAM scheme does not exceed the area required by conventional renaming mechanisms. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | L1-bandwidth aware thread allocation in multicore SMT processorsabstractImproving the utilization of shared resources is a key issue to increase performance in SMT processors. Recent work has focused on resource sharing policies to enhance the processor performance, but their proposals mainly concentrate on novel hardware mechanisms that adapt to the dynamic resource requirements of the running threads. This work addresses the L1 cache bandwidth problem in SMT processors experimentally on real hardware. Unlike previous work, this paper concentrates on thread allocation, by selecting the proper pair of co-runners to be launched to the same core. The relation between L1 bandwidth requirements of each benchmark and its performance (IPC) is analyzed. We found that for individual benchmarks, performance is strongly connected to L1 bandwidth consumption, and this observation remains valid when several co-runners are launched to the same SMT core. Based on these findings we propose two L1 bandwidth aware thread to core (t2c) allocation policies, namely Static and Dynamic t2c allocation, respectively. The aim of these policies is to properly balance L1 bandwidth requirements of the running threads among the processor cores. Experiments on a Xeon E5645 processor show that the proposed policies significantly improve the performance of the Linux OS kernel regardless the number of cores considered. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
PACT | 3 |
| 2013 | Combining RAM technologies for hard-error recovery in L1 data caches working at very-low power modesabstractLow-power modes in modern microprocessors rely on low frequencies and low voltages to reduce the energy budget. Nevertheless, manufacturing induced parameter variations can make SRAM cells unreliable producing hard errors at supply voltages below Vccmin. Vicente Lorente, Alejandro Valero, Julio Sahuquillo, Salvador Petit, Ramon Canal, Pedro López 0001, José Duato |
DATE | 4 |
| 2013 | Exploiting reuse information to reduce refresh energy in on-chip eDRAM cachesabstractThis work introduces a novel refresh mechanism that leverages reuse information to decide which blocks should be refreshed in an energy-aware eDRAM last-level cache. Experimental results show that, compared to a conventional eDRAM cache, the energy-aware approach achieves refresh energy savings up to 71%, while the reduction on the overall dynamic energy is by 65% with negligible performance losses. Alejandro Valero, Julio Sahuquillo, Salvador Petit, José Duato |
ICS | 3 |
| 2013 | Power-aware scheduling with effective task migration for real-time multicore embedded systemsabstractSUMMARY A major design issue in embedded systems is reducing the power consumption because batteries have a limited energy budget. For this purpose, several techniques such as dynamic voltage and frequency scaling (DVFS) or task migration are being used. DVFS allows reducing power by selecting the optimal voltage supply, whereas task migration achieves this effect by balancing the workload among cores. This paper focuses on power‐aware scheduling allowing task migration to reduce energy consumption in multicore embedded systems implementing DVFS capabilities. To address energy savings, the devised schedulers follow two main rules: migrations are allowed at specific points of time and only one task is allowed to migrate each time. Two algorithms have been proposed working under real‐time constraints. The simpler algorithm, namely, single option migration (SOM) only checks just one target core before performing a migration. In contrast, the multiple option migration (MOM) searches the optimal target core. In general, the MOM algorithm achieves better energy savings than the SOM algorithm, although differences are wider for a reduced number of cores and frequency/voltage levels. Moreover, the MOM algorithm reduces energy consumption as much as 40% over the worst fit algorithm. Copyright © 2012 John Wiley & Sons, Ltd. José Luis March, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Hardware-Based Generation of Independent Subtraces of Instructions in Clustered ProcessorsabstractMulticore chips are currently dominating the microprocessor market as designs that improve performance and sustain power consumption. However, complex core features must be still considered to provide good performance for existing sequential applications. An effective approach to reduce core complexity without dramatically sacrificing performance is to distribute critical processor structures by using clustered microarchitectures. In these designs, communication latency among clusters is a critical performance bottleneck, and a good steering algorithm is required to reduce intercluster communication. In this paper, we propose a new energy-efficient microarchitectural approach that reduces intercluster communication by detecting and generating independent chains of instructions, referred to as subtraces, from the execution of sequential programs. The devised mechanism has been modeled on an x86-based trace-cache processor, where subtraces are built in the fill unit, stored in a trace cache, and individually steered to different clusters. Experimental results show that the proposal reaches performance speedups around 7 and 15 percent for point-to-point and bus-based interconnects, respectively, while achieving energy savings of up to 12 percent. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 3 |
| 2012 | Analyzing the optimal ratio of SRAM banks in hybrid cachesabstractCache memories have been typically implemented with Static Random Access Memory (SRAM) technology. This technology presents a fast access time but high energy consumption and low density. As opposite, the recently appeared embedded Dynamic RAM (eDRAM) technology allows caches to be built with lower energy and area, although with a slower access time. The eDRAM technology provides important leakage and area savings, especially in huge Last-Level Caches (LLCs), which occupy almost half the silicon area in some recent microprocessors. This paper proposes a novel hybrid LLC, which combines SRAM and eDRAM banks to address the trade-off among performance, energy, and area. To this end, we explore the optimal percentage of SRAM and eDRAM banks that achieves the best target trade-off. Architectural mechanisms have been devised to keep the most likely accessed blocks in fast SRAM banks as well as to avoid unnecessary destructive reads. Experimental results show that, compared to a conventional SRAM LLC with the same storage capacity, performance degradation does not surpass, on average, 2.9% (even with 12.5% of banks built with SRAM technology), whereas area savings can be as high as 46% for a 1MB-16way LLC. For a 45nm technology node, the energy-delay squared product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is a quarter or an eighth of the cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ICCD | 3 |
| 2012 | Page-Based Memory Allocation Policies of Local and Remote Memory in Cluster ComputersabstractMain memory latencies have a strong impact on the overall execution time of the applications. The need of efficiently scheduling the costly DRAM memory resources in the different motherboards is a major concern in cluster computers. Most of these systems implement remote access capabilities which allow the OS to access to remote memory. In this context, efficient scheduling becomes even more critical since remote memory accesses may be several orders of magnitude higher than local accesses. These systems typically support interleaved memory at cache-block granularity. In contrast, in this paper we explore the impact on the system performance when allocating memory at OS page granularity. Experimental results show that simply supporting interleaved memory at OS page granularity is a feasible solution that does not impact on the performance of most of the benchmarks. Based on this observation we investigated the reasons of performance drops in those benchmarks showing unacceptable performance when working at page granularity. The results of this analysis lead us to propose two memory allocation policies, namely on-demand (OD) and Most-accessed in-local (Mail). The OD policy first places the requested pages in local memory, once this memory region is full, the subsequent memory pages are placed in remote memory. This policy shows good performance when the most accessed pages are requested and allocated before than the least accessed ones, which as proven in this work, is the most common case. This simple policy reaches performance improvements by 25% in some benchmarks with respect to a typical block interleaving memory system. Nevertheless, this strategy has poor performance when a noticeable amount of the least accessed pages are requested before than the most accessed ones. This performance drawback is solved by the Mail allocation policy by using profile information to guide the allocation of new pages. This scheme always outperforms the baseline block interleaving policy and, in some cases, improves the performance of the OD policy by 25%. Monica Serrano, Salvador Petit, Julio Sahuquillo, Rafael Ubal, Houcine Hassan, José Duato |
ICPADS | 2 |
| 2012 | Understanding Cache Hierarchy Contention in CMPs to Improve Job SchedulingabstractIn order to improve CMP performance, recent research has focused on scheduling to mitigate contention produced by the limited memory bandwidth. Nowadays, commercial CMPs implement multi-level cache hierarchies where last level caches are shared by at least two cache structures located at the immediately lower cache level. In turn, these caches can be shared by several multithreaded cores. In this microprocessor design, contention points may appear along the whole memory hierarchy. Moreover, this problem is expected to aggravate in future technologies, since the number of cores and hardware threads, and consequently the size of the shared caches increases with each microprocessor generation. In this paper we characterize the impact on performance of the different contention points that appear along the memory subsystem. Then, we propose a generic scheduling strategy for CMPs that takes into account the available bandwidth at each level of the cache hierarchy. The proposed strategy selects the processes to be co-scheduled and allocates them to cores in order to minimize contention effects. The proposal has been implemented and evaluated in a commercial single-threaded quad-core processor with a relatively small two-level cache hierarchy. Despite these potential contention limitations are less than in recent processor designs, compared to the Linux scheduler, the proposal reaches performance improvements up to 9% while these benefits (across the studied benchmark mixes) are always lower than 6% for a memory-aware scheduler that does not take into account the cache hierarchy. Moreover, in some cases the proposal doubles the speedup achieved by the memory-aware scheduler. Josué Feliu, Julio Sahuquillo, Salvador Petit, José Duato |
IPDPS | 3 |
| 2012 | Efficiently Handling Memory Accesses to Improve QoS in Multicore Systems under Real-Time ConstraintsabstractChip multiprocessors (CMPs) are becoming the common choice to implement embedded systems due to they achieve a good tradeoff between performance and power. Because of manufacturability reasons, CMPs use to implement one or several memory controllers, each one shared by a set of cores. Thus, memory requests from distinct cores compete among them when accessing to memory. This means that the memory access latency can widely vary depending on the co-runners and the memory controller scheduling policy, thus yielding to unpredictable behavior. This work focuses on the design of a memory controller to support workloads with real-time constraints, both hard real-time (HRT) and soft real-time (SRT) applications. These systems must guarantee the execution of HRT applications while improving the performance of the SRT applications. In this paper we propose two memory controller policies for multicore embedded systems: HR-first and ATR-first. The former prioritizes memory requests of HRT tasks, achieving important energy savings but poor performance for SRT applications. The latter gives priority to those HRT requests that are critical to guarantee schedulability. Results show that the ATR-first policy presents similar energy consumption as the HR-first policy while reducing the number of SRT deadline misses around 49%, on average, and reaching the fulfillment of all deadlines in some scenarios. José Luis March, Salvador Petit, Julio Sahuquillo, Houcine Hassan, José Duato |
SBAC-PAD | 2 |
| 2012 | Combining recency of information with selective random and a victim cache in last-level cachesabstractMemory latency has become an important performance bottleneck in current microprocessors. This problem aggravates as the number of cores sharing the same memory controller increases. To palliate this problem, a common solution is to implement cache hierarchies with large or huge Last-Level Cache (LLC) organizations. LLC memories are implemented with a high number of ways (e.g., 16) to reduce conflict misses. Typically, caches have implemented the LRU algorithm to exploit temporal locality, but its performance goes away from the optimal as the number of ways increases. In addition, the implementation of a strict LRU algorithm is costly in terms of area and power. This article focuses on a family of low-cost replacement strategies, whose implementation scales with the number of ways while maintaining the performance. The proposed strategies track the accessing order for just a few blocks, which cannot be replaced. The victim is randomly selected among those blocks exhibiting poor locality. Although, in general, the random policy helps improving the performance, in some applications the scheme fails with respect to the LRU policy leading to performance degradation. This drawback can be overcome by the addition of a small victim cache of the large LLC. Experimental results show that, using the best version of the family without victim cache, MPKI reduction falls in between 10% and 11% compared to a set of the most representative state-of-the-art algorithms, whereas the reduction grows up to 22% with respect to LRU. The proposal with victim cache achieves speedup improvements, on average, by 4% compared to LRU. In addition, it reduces dynamic energy, on average, up to 8%. Finally, compared to the studied algorithms, hardware complexity is largely reduced by the baseline algorithm of the family. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ACM Trans. Archit. Code Optim. | 3 |
| 2012 | Design, Performance, and Energy Consumption of eDRAM/SRAM Macrocells for L1 Data CachesabstractSRAM and DRAM have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology (eDRAM), thus overcoming the speed limit of typical DRAM cells. In this paper, we propose a hybrid n-bit macrocell that implements one SRAM cell and n-1 eDRAM cells. This cell is aimed at being used in an n-way set-associative first-level data cache. Architectural mechanisms (e.g., special writeback policies) have been devised to completely avoid refresh logic. Performance, energy, and area have been analyzed in detail. Experimental results show that using typical eDRAM capacitors, and compared to a conventional cache, a 4-way set-associative hybrid cache reduces both energy consumption and area up to 54 and 29 percent, respectively, while having negligible impact on performance (less than 2 percent). Alejandro Valero, Salvador Petit, Julio Sahuquillo, Pedro López 0001, José Duato |
IEEE Trans. Computers | 2 |
| 2012 | A cost-effective heuristic to schedule local and remote memory in cluster computers
Monica Serrano, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
J. Supercomput. | 3 |
| 2012 | A Sequentially Consistent Multiprocessor Architecture for Out-of-Order Retirement of InstructionsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Expanding the width of the instruction window can be highly beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory latencies. Based on the Validation Buffer (VB) architecture (a previously proposed out-of-order retirement, checkpoint-free architecture for single processors), this paper proposes a cost-effective, scalable, out-of-order retirement multiprocessor, capable of enforcing sequential consistency without impacting the design of the memory hierarchy or interconnect. Our simulation results indicate that utilizing a VB can speed up both relaxed and sequentially consistent in-order retirement in future multiprocessor systems by between 3 and 20 percent, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Impact on Performance and Energy of the Retention Time and Processor Frequency in L1 Macrocell-Based Data CachesabstractCache memories dissipate an important amount of the energy budget in current microprocessors. This is mainly due to cache cells are typically implemented with six transistors. To tackle this design concern, recent research has focused on the proposal of new cache cells. Ann-bit cache cell, namely macrocell, has been proposed in a previous work. This cell combines SRAM and eDRAM technologies with the aim of reducing energy consumption while maintaining the performance. The capacitance of eDRAM cells impacts on energy consumption and performance since these cells lose their state once the retention time expires. On such a case, data must be fetched from a lower level of the memory hierarchy, so negatively impacting on performance and energy consumption. As opposite, if the capacitance is too high, energy would be wasted without bringing performance benefits. This paper identifies the optimal capacitance for a given processor frequency. To this end, the tradeoff between performance and energy consumption of a macrocell-based cache has been evaluated varying the capacitance and frequency. Experimental results show that, compared to a conventional cache, performance losses are lower than 2% and energy savings are up to 55% for a cache with 10 fF capacitors and frequencies higher than 1 GHz. In addition, using trench capacitors, a 4-bit macrocell reduces by 29% the area of four conventional SRAM cells. Alejandro Valero, Julio Sahuquillo, Vicente Lorente, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Improving Last-Level Cache Performance by Exploiting the Concept of MRU-TourabstractLast-Level Caches (LLCs) implement the LRU algorithm to exploit temporal locality, but its performance is quite far of Belady's optimal algorithm as the number of ways increases. One of the main reasons because of LRU does not reach good performance in LLCs is that this policy forces a block to descend until the bottom of the stack before eviction. Nevertheless, most of the blocks that leave the MRU position are not referenced again before eviction. This work pursues to select candidate blocks to be victimized before reaching the bottom of the stack. To this end, this work defines the number of MRU-Tours (MRUTs) of a block as the number of times that a block enters in the MRU position during its live time. Based on the fact that most of the blocks exhibit a single MRUT, this work presents the family of MRUT-based algorithms aimed at exploiting this block behavior to improve performance. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 3 |
| 2011 | A Dynamic Power-Aware Partitioner with Task Migration for Multicore Embedded Systems
José Luis March, Julio Sahuquillo, Salvador Petit, Houcine Hassan, José Duato |
Euro-Par (1) | 3 |
| 2011 | A Cluster Computer Performance Predictor for Memory Scheduling
Monica Serrano, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICA3PP (2) | 4 |
| 2011 | MRU-Tour-based Replacement Algorithms for Last-Level CachesabstractMemory hierarchy design is a major concern in current microprocessors. Many research work focuses on the Last-Level Cache (LLC), which is designed to hide the long miss penalty of accessing to main memory. To reduce both capacity and conflict misses, LLCs are implemented as large memory structures with high associativities. To exploit temporal locality, LRU is the replacement algorithm usually implemented in caches. However, for a high-associative cache, its implementation is costly in terms of area and power consumption. Indeed, LRU is not well suited for the LLC, because as this cache level does not see all memory accesses, it cannot cope with temporal locality. In addition, blocks must descend down to the LRU position of the stack before eviction, even when they are not longer useful. In this paper, we show that most of the blocks are not referenced again once they leave the MRU position. Moreover, the probability of being referenced again does not depend on the location on the LRU stack. Based on these observations, we define the number of MRU-Tours (MRUTs) of a block as the number of times that a block occupies the MRU position while it is stored in the cache, and propose the MRUT replacement algorithm, which selects the block to be replaced among the blocks that show only one MRUT. Variations of this algorithm have been also proposed to exploit both MRUT behavior and recency of information. Experimental results show that, compared to LRU, the proposal reduces the MPKI up to 22%, while IPC is improved by 48%. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
SBAC-PAD | 3 |
| 2011 | A New Energy-Aware Dynamic Task Set Partitioning Algorithm for Soft and Hard Embedded Real-Time SystemsabstractPower consumption is a major design concern in current embedded systems. To deal with consumption, many systems apply dynamic voltage scaling (DVS) techniques which dynamically change the system speed depending on the workload characteristics. DVS costs in a multicore system can be reduced by sharing the same DVS regulator among the cores. In this context, to handle energy efficiently, the workload must be properly balanced among the cores. This paper proposes a new heuristic algorithm to balance the workload in an embedded system with a coarse-grain multithreaded multicore processor. This heuristic is aimed at improving the overlapping time between the memory and the processor while keeping balanced core utilizations. To this end, the heuristic dynamically drives the frequency/voltage level to guarantee deadline fulfillment of the hard real-time tasks as well as to achieve a good trade-off between deadline losses and energy savings of the soft real-time tasks. The proposed technique has been evaluated on a model of a contemporary high-end ARM embedded microprocessor executing a set of standard embedded benchmarks. Energy savings depend on the range of frequency/voltage levels that the DVS regulator implements. Experimental results show that with the proposed heuristic, when working with hard real-time tasks, the energy consumption is about 33% the energy dissipated by a system without DVS regulator and balancing heuristic. Moreover, when soft real-time tasks are also considered, the normalized consumption presents values ranging in between 8 and 70% depending on the scheduler aggressiveness. José Luis March, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
Comput. J. | 4 |
| 2010 | Exploiting subtrace-level parallelism in clustered processorsabstractNo abstract available. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 3 |
| 2010 | A Scheduling Heuristic to Handle Local and Remote Memory in Cluster ComputersabstractIn cluster computers, RAM memory is spread among the motherboards hosting the running applications. In these systems, it is common to constrain the memory address space of a given processor to the local motherboard. Constraining the system in this way is much cheaper than using a full-fledged shared memory implementation among motherboards. However, in this case, memory usage might widely differ among motherboards depending on the memory requirements of the applications running on each motherboard. In this context, if an application requires a huge quantity of RAM memory, the only feasible solution is to increase the amount of available memory in its local motherboard, even if the remaining ones are underused. Nevertheless, beyond a certain memory size, this memory budget increase becomes prohibitive. In this paper, we assume that the Remote Memory Access hardware used in a Hyper Transport based system allows applications to allocate the required memory from remote motherboards. We also analyze how the distribution of memory accesses among different memory locations (local or remote) impact on performance. Finally, an heuristic is devised to schedule local and remote memory among applications according to their requirements, and considering quality of service constraints. Monica Serrano, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
HPCC | 4 |
| 2010 | Extending a Multicore Multithread Simulator to Model Power-Aware Hard Real-Time Systems
José Luis March, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICA3PP (2) | 4 |
| 2010 | Out-of-order retirement of instructions in sequentially consistent multiprocessorsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Wide instruction windows are very beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory gap. In this paper, the Validation Buffer (VB) multiprocessor architecture is proposed as a cost-effective, checkpoint-free, scalable approach to retire instructions out of program order, while still enforcing sequential consistency, and without impacting the memory hierarchy or interconnect. Experimental results show that utilizing the Validation Buffer can speed up both release and sequentially consistent in-order retirement in future multiprocessor systems by between 3% and 20%, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
ICCD | 3 |
| 2010 | Balancing Task Resource Requirements in Embedded Multithreaded Multicore Processors to Reduce Power ConsumptionabstractPower consumption is a major design issue in modern microprocessors. Hence, power reduction techniques, like Dynamic Voltage Scaling (DVS), are being widely implemented. Unfortunately, they impact on the task execution time so difficulting schedulability of hard real-time applications. To deal with this problem, this paper proposes a power-aware scheduler for coarse-grain embedded multicore processors implementing global DVS. To this end, this work presents two heuristics, namely Balanced Memory and Balanced CPU, which distribute the task set among cores focusing on resource utilization. Results show that with respect to a system not implementing DVS, two or five DVS levels achieve energy savings by about 35% or 51%, respectively. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
PDP | 4 |
| 2009 | An Efficient Low-Complexity Alternative to the ROB for Out-of-Order Retirement of InstructionsabstractCurrent superscalar processors use a reorder buffer (ROB) to support speculation, precise exceptions, and register reclamation. Instructions are retired from this structure in program order, which may lead to significant performance degradation if a long latency operation blocks the ROB head. In this paper, a checkpoint-free out-of-order commit architecture is proposed, which replaces the ROB with a small structure called validation buffer (VB) from which instructions are retired as soon as their speculative state is resolved. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. Experimental results show that the VB microarchitecture is much more efficient than a ROB-based microprocessor. For example, a 32-entry VB provides similar performance to a 256-entry ROB, while reducing the utilization of other major processor structures. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001, José Duato |
DSD | 1 |
| 2009 | Paired ROBs: A Cost-Effective Reorder Buffer Sharing Strategy for SMT Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 3 |
| 2009 | A power-aware hybrid RAM-CAM renaming mechanism for fast recoveryabstractModern superscalar processors implement register renaming by using either RAM or CAM tables. The design of these structures should address their access time and misprediction recovery penalty. While direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. Although they are more complex and slower, CAMs usually match the processor cycle in current designs. However, they do not scale with the number of physical registers and the pipeline width. In this paper we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides the current mappings quickly; on mispeculation, a low-complexity CAM enables immediate recovery and further register renaming. Compared to an ideal CAM in a 4-way state-of-the-art superscalar microprocessor, and for almost the same performance (1% slowdown) and area (95% of the ideal CAM size), the proposed scheme consumes about 90% less dynamic energy. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
ICCD | 1 |
| 2009 | Dynamic task set partitioning based on balancing memory requirements to reduce power consumptionabstractBecause of technology advances power consumption has emerged up as an important design issue in modern high-performance microprocessors. As a consequence, research on reducing power consumption has become a hot research topic. Different ways to reduce power consumption consist on using processors that do not implement the most power-hungry microarchitectural mechanisms, attacking hot spots, or reducing consumption in the larger microprocessor components like the cache. Unlike these works which focus on specific parts of the microprocessor, Dynamic Voltage Scaling (DVS) is a technique which applies on the whole microprocessor die. This technique allows the system to work at different frequency/voltage levels. DVS costs in a multicore system can be reduced by sharing the same DVS regulator among the cores (global DVS). In this context, to handle energy efficiently, the workload must be properly balanced among the cores. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
ICS | 4 |
| 2009 | An hybrid eDRAM/SRAM macrocell to implement first-level data cachesabstractSRAM and DRAM cells have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology, thus overcoming the speed limit of typical DRAM cells. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Vicente Lorente, Ramon Canal, Pedro López 0001, José Duato |
MICRO | 3 |
| 2009 | A Complexity-Effective Out-of-Order Retirement MicroarchitectureabstractCurrent superscalar processors commit instructions in program order by using a reorder buffer (ROB). The ROB provides support for speculation, precise exceptions, and register reclamation. However, committing instructions in program order may lead to significant performance degradation if a long latency operation blocks the ROB head. Several proposals have been published to deal with this problem. Most of them retire instructions speculatively. However, as speculation may fail, checkpoints are required in order to rollback the processor to a precise state, which requires both extra hardware to manage checkpoints and the enlargement of other major processor structures, which, in turn, might impact the processor cycle. This paper focuses on out-of-order commit in a nonspeculative way, thus, avoiding checkpointing. To this end, we replace the ROB with a validation buffer (VB) structure. This structure keeps dispatched instructions until they are nonspeculative or mispeculated, which allows an early retirement. By doing so, the performance bottleneck is largely alleviated. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. As experimental results show, the VB structure is much more efficient than a typical ROB since, with only 32 entries, it achieves a performance close to an in-order commit microprocessor using a 256-entry ROB. Salvador Petit, Julio Sahuquillo, Pedro López 0001, Rafael Ubal, José Duato |
IEEE Trans. Computers | 1 |
| 2008 | Reducing the Number of Bits in the BTB to Attack the Branch Predictor Hot-Spot
Noel Tomás, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 3 |
| 2008 | A simple power-aware scheduling for multicore systems when running real-time applicationsabstractHigh-performance microprocessors, e.g., multithreaded and multicore processors, are being implemented in embedded real-time systems because of the increasing computational requirements. These complex microprocessors have two major drawbacks when they are used for real-time purposes. First, their complexity difficults the calculation of the WCET (worst case execution time). Second, power consumption requirements are much larger, which is a major concern in these systems. In this paper we propose a novel soft power-aware real-time scheduler for a state-of-the-art multicore multithreaded processor, which implements dynamic voltage scaling techniques. The proposed scheduler reduces the energy consumption while satisfying the constraints of soft real-time applications. Different scheduling alternatives have been evaluated, and experimental results show that using a fair scheduling policy, the proposed algorithm provides, on average, energy savings ranging from 34% to 74%. Diana Bautista, Julio Sahuquillo, Houcine Hassan, Salvador Petit, José Duato |
IPDPS | 4 |
| 2008 | The impact of out-of-order commit in coarse-grain, fine-grain and simultaneous multithreaded architecturesabstractMultithreaded processors in their different organizations (simultaneous, coarse grain and fine grain) have been shown as effective architectures to reduce the issue waste. On the other hand, retiring instructions from the pipeline in an out-of-order fashion helps to unclog the ROB when a long latency instruction reaches its head. This further contributes to maintain a higher utilization of the available issue bandwidth. In this paper, we evaluate the impact of retiring instructions out of order on different multithreaded architectures and different instruction fetch policies, using the recently proposed Validation Buffer microarchitecture as baseline out-of-order commit technique. Experimental results show that, for the same performance, out-of-order commit permits to reduce multithread hardware complexity (e.g., fine grain multithreading with a lower number of supported threads). Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IPDPS | 3 |
| 2007 | VB-MT: Design Issues and Performance of the Validation Buffer Microarchitecture for Multithreaded Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 3 |
| 2007 | Multi2Sim: A Simulation Framework to Evaluate Multicore-Multithreaded ProcessorsabstractCurrent microprocessors are based in complex designs, integrating different components on a single chip, such as hardware threads, processor cores, memory hierarchy or interconnection networks. The permanent need of evaluating new designs on each of these components motivates the development of tools which simulate the system working as a whole. In this paper, we present the Multi2Sim simulation framework, which models the major components of incoming systems, and is intended to cover the limitations of existing simulators. A set of simulation examples is also included for illustrative purposes. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 3 |
| 2006 | Applying the zeros switch-off technique to reduce static energy in data cachesabstractZeros switch-off is a leakage energy reduction technique applicable to cache memories. It works at the cache word level by removing the power supply of all or part of its most significant bytes when they store a zero, taking advantage of the high percentage of zero data bits in common programs. Experimental results, obtained by using the SPEC2000 benchmarks suite, show that the average leakage energy savings reach 60.3% with no IPC loss indeed. The proposed technique can be combined with other existing energy reduction techniques, reaching, on average, 67.3% savings with 0.5% IPC losses Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 3 |
| 2006 | Addressing a workload characterization study to the design of consistency protocols
Salvador Petit, Julio Sahuquillo, Ana Pont, David R. Kaeli |
J. Supercomput. | 1 |
| 2005 | Exploring the performance of split data cache schemes on superscalar processors and symmetric multiprocessors
Julio Sahuquillo, Salvador Petit, Ana Pont, Veljko M. Milutinovic |
J. Syst. Archit. | 2 |
| 2004 | Characterizing the Dynamic Behavior of Workload Execution in SVM systemsabstractThe overhead associated with software management of shared virtual memory (SVM) systems can seriously impact overall system performance. One way to remedy this situation is to design more efficient SVM consistency protocols. In this paper we study a number of parallel workload characteristics that can negatively impact the performance of SVM systems. We attempt to quantify the sources of performance loss in some parallel workloads. Our goal is to better understand these characteristics, enabling us to develop SVM protocols that can adjust to dynamics in workload behavior. This paper has three main contributions: i) we measure the contention for synchronization resources, showing how applications exhibit distinct phases during their execution, ii) we quantify the relationship between page size and fragmentation/false sharing while varying the sharing unit size, and iii) we study the synergies between the contention for synchronization resources and fragmentation/false sharing, providing hints for developing improved protocols. Salvador Petit, Julio Sahuquillo, Ana Pont, David R. Kaeli |
SBAC-PAD | 1 |
| 2001 | About the sensitivity of the HLRC-DU protocol on diff and page sizesabstractRecent research on software distributed shared memory systems has focused on consistency protocols for improving performance. Home Lazy Release Consistency (HLRC) protocols have been widely adopted due to their performance advantages. Usually, these protocols invalidate pages through write notices. Variants of these protocols propose some criterion to update data of the corresponding pages instead of invalidating. In a previous paper, we proposed the HLRC-DU protocol, which is an improved version of the HLRC protocol. The HLRC-DU embeds update information in those write notices whose corresponding diff size is less than a given threshold, invalidating the remainder. The threshold trades off network bandwidth with update perfonnance. In this paper, we study the HLRC-DUprotocol’s sensitivity to page size and the threshold size selection. Our results show that while the page size slightly impacts performance, that our protocols are highly sensitive to the threshold value. Salvador Petit, Julio Sahuquillo, Ana Pont |
ISPASS | 1 |