EDBT 2026 Demo / reviewers in the wild / expert
Juan M. Cebrian
dblp:44/7979 · also Juan Manuel Cebrian Gonzalez
· DBLP profile ↗
33ranked-venue papers
15as first author
8since 2021 · last 2025
0000-0002-3731-9301ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 13 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Precise characterization of coherence activity in multicores using gem5
Joaquín Ferrer, Juan M. Cebrian, Ricardo Fernández-Pascual, Manuel E. Acacio |
J. Supercomput. | 2 |
| 2024 | Bounding Speculative Execution of Atomic Regions to a Single RetryabstractMutual exclusion has long served as a fundamental construct in parallel programs. Despite a long history of optimizing the lower-level lock and unlock operations used to enforce mutual exclusion, such operations largely dictate performance in parallel programs. Speculative Lock Elision, and more generally Hardware Transactional Memory, allow executing atomic regions (ARs) concurrently and speculatively, and ensure correctness by using conflict detection. However, practical implementations of these ideas are best-effort and, in case of conflicts, the execution of ARs is retried a predetermined number of times before falling back to mutual exclusion. Eduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto Ros 0001 |
ASPLOS (4) | 2 |
| 2024 | Hardware Cache Locking for All Memory UpdatesabstractMany applications need to perform operations that involve reading a value from memory, modifying it, and then writing it back. Multiple architectures provide hardware support for these operations via read-modify-write (RMW) instructions. The primary benefit is that the read can request a cacheline with write permissions, reducing coherence protocol overhead since the write will find the cacheline with appropriate permissions. RMWs can be either atomic or non-atomic. Atomic RMWs, used for synchronization, commonly require (i) locking the cacheline to guarantee atomicity by preventing invalidations and (ii) enforcing serialization of instructions in the program (e.g., via memory fences), which may cause performance degradation based on the implemented memory consistency model. Non-atomic RMWs, while not requiring such strict measures, should only be used in data-race free code sections. However, other cores may invalidate a cacheline during a non-atomic RMW (e.g., due to false sharing), flushing the pipeline and causing the loss of write permissions obtained by the read, which is detrimental to performance. In this work, we propose a microarchitectural mechanism that enables non-atomic RMWs to fetch the cacheline locking it, thus preventing other cores from “stealing” the cacheline while allowing them to run concurrently with other instructions in the same core. Our proposal enables concurrent hardware cache locking for multiple non-atomic RMW s while guaranteeing deadlock freedom and no programmer/compiler intervention. We also propose a lock-chaining mechanism to allow multiple consecutive memory updates to the same cacheline up to a predefined maximum (to prevent starvation and load imbalance). Our evaluation using gem5 full-system simulator shows that for an eight-core configuration, our proposal improves performance by up to 5.36 % (2.05 % on average), requiring just 45 bytes of storage per core. Ashkan Asgharzadeh, Eduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto Ros 0001 |
ICCD | 3 |
| 2024 | Temporarily Unauthorized Stores: Write First, Ask for Permission Laterabstractx86 processors implement a total store order (x86-TSO) consistency model, which requires stores to update memory in a sequenced manner. The latency of stores is then hidden by the store buffer (SB), which holds stores until the write is performed. On a long latency cache miss, however, stores block the SB, eventually stalling the processor and degrading performance. Contemporary industrial high-performance processors deal with this situation by overprovisioning the size of the SB, but this comes at the cost of energy and latency overheads. In this work, we remove the stalls caused by stores blocked at the head of the SB while reusing existing processor resources, either improving performance when SB size is kept constant or maintaining performance while reducing SB size. Our proposal, Temporarily Unauthorized Stores (TUS), achieves this by extending the functionality of 1) the write combining buffers, to allow them to coalesce stores while maintaining x86- TSO consistency, and 2) immediately write data to the first-level cache upon a miss (i.e., providing an always-hit illusion) but temporarily keeping the written data invisible to the cache coherence protocol, i.e., these stores are temporarily unauthorized. TUS makes temporarily unauthorized stores visible in x86- TSO order without speculation or rollbacks once write permission is obtained. In essence, TUS logically transforms the write combining buffers and the first-level cache into an “extension” of the SB. TUS improves performance by up to 26 % (3.2 % on average) while reducing the total energy-delay-product (EDP) by up to 35.9% (6.4% on average) for SB-bound benchmarks with a 114-entry SB compared to our baseline architecture with an SB of the same size. When configured with a 32-entry SB, TUS yields a performance improvement of 2 % over a 114-entry SB baseline while reducing SB energy per search by a factor of 2 x, SB area by 21 %, and store-to-Ioad forwarding latency from 5 to 3 cycles. Juan M. Cebrian, Magnus Jahre, Alberto Ros 0001 |
MICRO | 1 |
| 2023 | Near-optimal multi-accelerator architectures for predictive maintenance at the edge
Mostafa Koraei, Juan M. Cebrian, Magnus Jahre |
Future Gener. Comput. Syst. | 2 |
| 2022 | Free atomics: hardware atomic operations without fencesabstractAtomic Read-Modify-Write (RMW) instructions are primitive synchronization operations implemented in hardware that provide the building blocks for higher-abstraction synchronization mechanisms to programmers. According to publicly available documentation, current x86 implementations serialize atomic RMW operations, i.e., the store buffer is drained before issuing atomic RMWs and subsequent memory operations are stalled until the atomic RMW commits. This serialization, carried out by memory fences, incurs a performance cost which is expected to increase with deeper pipelines. Ashkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras, Alberto Ros 0001 |
ISCA | 2 |
| 2022 | Compiler-Assisted Compaction/Restoration of SIMD InstructionsabstractVector processors (e.g., SIMD or GPUs) are ubiquitous in high performance systems. All the supercomputers in the world exploit data-level parallelism (DLP), for example by using single instructions to operate over several data elements. Improving vector processing is therefore key for exascale computing. However, despite its potential, vector code generation and execution have significant challenges. Among these challenges, control flow divergence is one of the main performance limiting factors. Most modern vector instruction sets, including SIMD, rely on predication to support divergence control. Nevertheless, the performance and energy consumption in predicated codes is usually insensitive to the number of active elements in a predicated mask. Since the trend is that vector register size increases, the energy efficiency of exascale computing systems will become sub-optimal. This article proposes a novel approach to improve execution efficiency in predicated vector codes, the Compiler-Assisted Compaction/Restoration (CACR) technique. Baseline CR delays predicated SIMD instructions with inactive elements, compacting active elements from instances of the same instruction of consecutive loop iterations. Compacted elements form an equivalentdensevector instruction. After executing the dense instructions, their results are restored to the original instructions. However, CR has a significant performance and energy penalty when it fails to find active elements, either due to lack of resources when unrolling or because of inter-loop dependencies. In CACR, the compiler analyzes the code looking for key information required to configure CR. Then, it passes this information to the processor via new instructions inserted in the code. This prevents CR from waiting for active elements on scenarios when it would fail to form dense instructions. Simulated results (gem5) show that CACR improves performance by up to 29 percent and reduces dynamic energy by up to 24.2 percent on average, for a a set of applications with predicated execution. The baseline CR only achieves 18.6 percent performance and 14 percent energy improvements for the same configuration and applications. Juan M. Cebrian, Thibaud Balem, Adrián Barredo, Marc Casas, Miquel Moretó, Alberto Ros 0001, Alexandra Jimborean |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Efficient, Distributed, and Non-Speculative Multi-Address Atomic OperationsabstractCritical sections that read, modify, and write (RMW) a small set of addresses are common in parallel applications and concurrent data structures. However, to escape from the intricacies of fine-grained locks, which require reasoning about all possible thread interleavings, programmers often resort to coarse-grained locks to ensure atomicity. This results in atomic protection of a much larger set of potentially conflicting addresses, and, consequently, increased lock contention and unneeded serialization. As many before us have observed, these problems would be solved if only general RMW multi-address atomic operations were available, but current proposals are impractical because of deadlock scenarios that appear due to resource limitations. Alternatively, transactional memory can detect conflicts at run-time aiming to maximize concurrency, but it has significant overheads in highly-contended critical sections. Eduardo José Gómez-Hernández, Juan M. Cebrian, J. Rubén Titos Gil, Stefanos Kaxiras, Alberto Ros 0001 |
MICRO | 2 |
| 2020 | Improving Predication Efficiency through Compaction/Restoration of SIMD InstructionsabstractVector processors offer a wide range of unexplored opportunities to improve performance and energy efficiency. However, despite its potential, vector code generation and execution have significant challenges, the most relevant ones being control flow divergence. Most modern processors including SIMD extensions (such as AVX) rely on predication to support divergence control. In predicated codes, performance and energy consumption are usually insensitive to the number of true values in a predicated mask. This implies that the system efficiency becomes sub-optimal as vector length increases. In this paper we focus on SIMD extensions and propose a novel approach to improve execution efficiency in predicated SIMD instructions, the Compaction/Restoration (CR) technique. CR delays predicated SIMD instructions with inactive elements and compacts them with instances of the same instruction from different loop iterations to form an equivalent dense vector instruction, where, in the best case, all the elements are active. After executing such dense instructions, their results are restored to the original instructions. Our evaluation shows that CR improves performance by up to 25% and reduces dynamic energy consumption by up to 43% on real unmodified applications with predicated execution. Moreover, CR allows executing unmodified legacy code with short vector instructions (AVX-2) on newer architectures with wider vectors (AVX-512), achieving up to 56% performance benefits. Adrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas, Mateo Valero |
HPCA | 2 |
| 2020 | Boosting Store Buffer Efficiency with Store-Prefetch BurstsabstractVirtually all processors today employ a store buffer (SB) to hide store latency. However, when the store buffer is full, store latency is exposed to the processor causing pipeline stalls. The default strategies to mitigate these stalls are to issue prefetch for ownership requests when store instructions commit and to continuously increase the store buffer size. While these strategies considerably increase memory-level parallelism for stores, there are still applications that suffer deeply from stalls caused by the store buffer. Even worse, store-buffer induced stalls increase considerably when simultaneous multi-threading is enabled, as the store buffer is statically partitioned among the threads.In this paper, we propose a highly selective and very aggressive prefetching strategy to minimize store-buffer induced stalls. Our proposal, Store-Prefetch Burst (SPB), is based on the following insights: i) the majority of store-buffer induced stalls are caused by a few stores; ii) the access pattern of such stores are easily predictable; and iii) the latency of the stores is not commonly hidden by standard cache prefetchers, as hiding their latency would require tremendous prefetch aggressiveness. SPB accurately detects contiguous store-access patterns (requiring just 67 bits of storage) and prefetches the remaining memory blocks of the accessed page in a single burst request to the L1 controller. SPB matches the performance of a 1024-entry SB implementation on a 56-entry SB (i.e., Skylake architecture). For a 14-entry SB (e.g., running four logical cores), it achieves 95.0% of that ideal performance, on average, for SPEC CPU 2017. Additionally, a 20-entry store buffer that incorporates SPB achieves the average performance of a standard 56-entry store buffer. Juan M. Cebrian, Stefanos Kaxiras, Alberto Ros 0001 |
MICRO | 1 |
| 2020 | Semi-automatic validation of cycle-accurate simulation infrastructures: The case for gem5-x86
Juan M. Cebrian, Adrián Barredo, Helena Caminal, Miquel Moretó, Marc Casas, Mateo Valero |
Future Gener. Comput. Syst. | 1 |
| 2020 | High-throughput fuzzy clustering on heterogeneous architectures
Juan M. Cebrian, Baldomero Imbernon, Jesús A. Soto, José M. García 0001, José M. Cecilia |
Future Gener. Comput. Syst. | 1 |
| 2020 | Using Arm's scalable vector extension on stencil codes
Adrià Armejach, Helena Caminal, Juan M. Cebrian, Rubén Langarita, Rekai González-Alberquilla, Chris Adeniyi-Jones, Mateo Valero, Marc Casas, Miquel Moretó |
J. Supercomput. | 3 |
| 2020 | Efficiency analysis of modern vector architectures: vector ALU sizes, core counts and clock frequencies
Adrián Barredo, Juan M. Cebrian, Mateo Valero, Marc Casas, Miquel Moretó |
J. Supercomput. | 2 |
| 2020 | Scalability analysis of AVX-512 extensions
Juan M. Cebrian, Lasse Natvig, Magnus Jahre |
J. Supercomput. | 1 |
| 2019 | POSTER: An Optimized Predication Execution for SIMD ExtensionsabstractVector processing is a widely used technique to improve performance and energy efficiency in modern processors. Most of them rely on predication to support divergence control. However, performance and energy consumption in predicated instructions are usually independent on the number of true values in a mask. This means that the efficiency of the system becomes sub-optimal as vector length increases. In this work we propose the Optimized Predication Execution (OPE) technique. OPE delays the execution of sparse masked vector instructions sharing the same PC, extracts their active elements and creates a new dense instruction with a higher mask density. After executing such dense instruction, results are restored to the original sparse instructions. Our approach improves performance by up to 25% and reduces dynamic energy consumption by up to 43% on real applications with predication. Adrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas, Mateo Valero |
PACT | 2 |
| 2018 | Stencil codes on a vector length agnostic architectureabstractData-level parallelism is frequently ignored or underutilized. Achieved through vector/SIMD capabilities, it can provide substantial performance improvements on top of widely used techniques such as thread-level parallelism. However, manual vectorization is a tedious and costly process that needs to be repeated for each specific instruction set or register size. In addition, automatic compiler vectorization is susceptible to code complexity, and usually limited due to data and control dependencies. To address some these issues, Arm recently released a new vector ISA, the Scalable Vector Extension (SVE), which is Vector-Length Agnostic (VLA). VLA enables the generation of binary files that run regardless of the physical vector register length. Adrià Armejach, Helena Caminal, Juan M. Cebrian, Rekai González-Alberquilla, Chris Adeniyi-Jones, Mateo Valero, Marc Casas, Miquel Moretó |
PACT | 3 |
| 2018 | Performance and energy effects on task-based parallelized applications - User-directed versus manual vectorization
Helena Caminal, Diego Caballero, Juan M. Cebrian, Roger Ferrer, Marc Casas, Miquel Moretó, Xavier Martorell, Mateo Valero |
J. Supercomput. | 3 |
| 2018 | A vectorized k-means algorithm for compressed datasets: design and experimental analysis
Abdullah Al Hasib, Juan M. Cebrian, Lasse Natvig |
J. Supercomput. | 2 |
| 2017 | A dedicated private-shared cache design for scalable multiprocessorsabstractSummary Most modern architectures are based on a shared‐memory design. Correctness of these architectures is ensured by means of coherence protocols and consistency models. However, performance and scalability of shared‐memory systems is usually constrained by the amount and size of the messages used to keep the memory subsystem coherent. This is not only important in high performance computing, but also in low power embedded systems, specially if coherence is required between different components of the system‐on‐chip. We argue that using the same mechanism to keep coherence for all memory accesses can be counterproductive, because it incurs unnecessary overhead for data addresses that would remain coherent after the access (i.e., private data and read‐only shared data). This paper proposes the use of dedicated caches for two different kinds of data (i) data that can be accessed without contacting other nodes and (ii) modifiable shared data. The private cache (L1P) will be independent for each core and will store private data and read‐only shared data. On the other hand, the shared cache (L1S), will be logically shared but physically distributed for all cores. With this design, we can significantly simplify the coherence protocol, reduce the on‐chip area requirements and reduce invalidation time. However, this dedicated cache design requires a classification mechanism to detect the nature of the data that is being accessed. Results show two drawbacks to this approach: first, the accuracy of the classification mechanism has a huge impact on performance. Second, a traditional interconnection network is not optimal for accessing the L1S, increasing register‐to‐cache latency when accessing shared data. Copyright © 2016 John Wiley & Sons, Ltd. Juan M. Cebrian, Ricardo Fernández-Pascual, Alexandra Jimborean, Manuel E. Acacio, Alberto Ros 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Transient Temperature Prediction for Aging Thermal Sensors Using Artificial Neural NetworkabstractAs technology scales down and power density increases, the temperature sensor characteristics will drift, leading to temperature errors which increase over time. Transistor aging is one of the leading contributors to temperature sensing inaccuracies. The prominent aging failure mechanisms like Negative Bias Temperature Instability (NBTI), Hot Carrier Injection (HCI) and electromigration have emerged as the main sources of system unreliability which manifest as an increase in the propagation delay over time. On-chip thermal sensors are not immune to this phenomenon and get affected by these aging mechanisms. Thermal sensor aging exacerbated by increased temperatures leads to temperature sensing inaccuracies requiring repeated sensor calibration. In this work, we propose a novel approach of using performance metrics to predict the transient temperature profile of an application as seen by the aging thermal sensor. Firstly, we make offline profiling of applications and then cluster them into groups using k-means clustering mechanism. Then we use a neural network model to predict the thermal profile of a new application given its performance metrics. The forecasting ability of our model is accessed using MSE and RMSE. This approach is highly scalable and can be used to predict future temperatures which can then be used for run-time dynamic thermal management of multi-core systems. Kameswar Rao Vaddina, Juan M. Cebrian, Lasse Natvig |
PDP | 2 |
| 2015 | Early Experiences with Separate Caches for Private and Shared DataabstractShared-memory architectures have become predominant in modern multi-core microprocessors in all market segments, from embedded to high performance computing. Correctness of these architectures is ensured by means of coherence protocols and consistency models. Performance and scalability of shared-memory systems is usually limited by the amount and size of the messages used to keep the memory subsystem coherent. Moreover, we believe that blindly keeping coherence for all memory accesses can be counterproductive, since it incurs in unnecessary overhead for data that will remain coherent after the access. Having this in mind, in this paper we propose the use of dedicated caches for private (+shared read-only) and shared data. The private cache (L1P) will be independent for each core while the shared cache (L1S) will be logically shared but physically distributed for all cores. This separation should allow us to simplify the coherence protocol, reduce the on-chip area requirements and reduce invalidation time with minimal impact on performance. The dedicated cache design requires a classification mechanism to detect private and shared data. In our evaluation we will use a classification mechanism that operates at the operating system (OS) level (page granularity). Results show two drawbacks to this approach: first, the selected classification mechanism has too many false positives, thus becoming an important limiting factor. Second, a traditional interconnection network is not optimal for accessing the L1S, and a custom network design is needed. These drawbacks lead to important performance degradation due to the additional latency when accessing the shared data. Juan M. Cebrian, Alberto Ros 0001, Ricardo Fernández-Pascual, Manuel E. Acacio |
e-Science | 1 |
| 2015 | V-PFORDelta: Data Compression for Energy Efficient Computation of Time SeriesabstractChip multiprocessors (CMPs) and heterogeneous architectures have become predominant in all market segments, from embedded to high performance computing. These architectures exacerbate on-chip data requirements, creating additional pressure on the memory subsystem. Consequently, efficient utilization of on-chip memory space becomes critical for data intensive applications. A promising means of addressing this challenge is to use an effective compression method to reduce the data transmitted along the memory hierarchy. In this paper we present V-PFORDelta, a real-time vectorized integer differential compression method for memory bound applications. We evaluate the effectiveness of our SIMD (Single Instruction Multiple Data stream) based compression method on an industrial hydrological time series data processing kernel. We analyzed both Streaming SIMD Extensions (SSE) and Advanced Vector Extensions 2 (AVX2) versions of the compression method. Results show that the performance and energy efficiency can be improved up to a factor of 3.1 and 8.2, respectively. The proposed method not only outperforms the uncompressed SIMD implementations of the hydrological kernel, but also reduces the data storage requirements by a factor of 1.56x to 3.38x, depending on the analyzed dataset. Abdullah Al Hasib, Juan M. Cebrian, Lasse Natvig |
HiPC | 2 |
| 2015 | Soft-error mitigation by means of decoupled transactional memory threads
Daniel Sánchez 0004, Juan M. Cebrian, José M. García 0001, Juan L. Aragón |
Distributed Comput. | 2 |
| 2014 | Optimized hardware for suboptimal software: The case for SIMD-aware benchmarksabstractEvaluation of new architectural proposals against real applications is a necessary step in academic research. However, providing benchmarks that keep up with new architectural changes has become a real challenge. If benchmarks don't cover the most common architectural features, architects may end up under/over estimating the impact of their contributions. In this work, we extend the PARSEC benchmark suite with SIMD capabilities to provide an enhanced evaluation framework for new academic/industry proposals. We then perform a detailed energy and performance evaluation of this commonly used application set on different platforms (Intel®and ARM®processors). Our results show how SIMD code alters scalability, energy efficiency and hardware requirements. Performance and energy efficiency improvements depend greatly on the fraction of code that we can actually vectorize (up to 50×). Our enhancements are based in a custom built wrapper library compatible with SSE, AVX and NEON to facilitate general vectorization. We aim to distribute the source code to reinforce the evaluation process of new proposals for computing systems. Juan M. Cebrian, Magnus Jahre, Lasse Natvig |
ISPASS | 1 |
| 2014 | Toward energy efficiency in heterogeneous processors: findings on virtual screening methodsabstractABSTRACT The integration of the latest breakthroughs in computational modeling and high performance computing (HPC) has leveraged advances in the fields of healthcare and drug discovery, among others. By integrating all these developments together, scientists are creating new exciting personal therapeutic strategies for living longer that were unimaginable not that long ago. However, we are witnessing the biggest revolution in HPC in the last decade. Several graphics processing unit architectures have established their niche in the HPC arena but at the expense of an excessive power and heat. A solution for this important problem is based on heterogeneity. In this paper, we analyze power consumption on heterogeneous systems, benchmarking a bioinformatics kernel within the framework of virtual screening methods. Cores and frequencies are tuned to further improve the performance or energy efficiency on those architectures. Our experimental results show that targeted low‐cost systems are the lowest power consumption platforms, although the most energy efficient platform and the best suited for performance improvement is the Kepler GK110 graphics processing unit from Nvidia by using compute unified device architecture. Finally, the open computing language version of virtual screening shows a remarkable performance penalty compared with its compute unified device architecture counterpart. Copyright © 2013 John Wiley & Sons, Ltd. Ginés D. Guerrero, Juan M. Cebrian, Horacio Emilio Pérez Sánchez, José M. García 0001, Manuel Ujaldon, José M. Cecilia |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Managing power constraints in a single-core scenario through power tokens
Juan M. Cebrian, Daniel Sánchez 0004, Juan L. Aragón, Stefanos Kaxiras |
J. Supercomput. | 1 |
| 2013 | Modeling the impact of permanent faults in cachesabstractThe traditional performance cost benefits we have enjoyed for decades from technology scaling are challenged by several critical constraints including reliability. Increases in static and dynamic variations are leading to higher probability of parametric and wear-out failures and are elevating reliability into a prime design constraint. In particular, SRAM cells used to build caches that dominate the processor area are usually minimum sized and more prone to failure. It is therefore of paramount importance to develop effective methodologies that facilitate the exploration of reliability techniques for caches. To this end, we present an analytical model that can determine for a given cache configuration, address trace, and random probability of permanent cell failure the exact expected miss rate and its standard deviation when blocks with faulty bits are disabled. What distinguishes our model is that it is fully analytical, it avoids the use of fault maps, and yet, it is both exact and simpler than previous approaches. The analytical model is used to produce the miss-rate trends ( expected miss-rate ) for future technology nodes for both uncorrelated and clustered faults. Some of the key findings based on the proposed model are (i) block disabling has a negligible impact on the expected miss-rate unless probability of failure is equal or greater than 2.6e-4, (ii) the fault map methodology can accurately calculate the expected miss-rate as long as 1,000 to 10,000 fault maps are used, and (iii) the expected miss-rate for execution of parallel applications increases with the number of threads and is more pronounced for a given probability of failure as compared to sequential execution. Daniel Sánchez 0004, Yiannakis Sazeides, Juan M. Cebrian, José M. García 0001, Juan L. Aragón |
ACM Trans. Archit. Code Optim. | 3 |
| 2011 | Token3D: Reducing Temperature in 3D Die-Stacked CMPs through Cycle-Level Power Control Mechanisms
Juan M. Cebrian, Juan L. Aragón, Stefanos Kaxiras |
Euro-Par (1) | 1 |
| 2011 | Power Token Balancing: Adapting CMPs to Power Constraints for Parallel Multithreaded WorkloadsabstractIn the recent years virtually all processor architectures employ multiple cores per chip (CMPs). It is possible to use legacy (i.e., single-core) power saving techniques in CMPs which run either sequential applications or independent multithreaded workloads. However, new challenges arise when running parallel shared-memory applications. In the later case, sacrificing some performance in a single core (thread) in order to be more energy-efficient might unintentionally delay the rest of cores (threads) due to synchronization points (locks/barriers), therefore, harming the performance of the whole application. CMPs increasingly face thermal and power-related problems during their typical use. Such problems can be solved by setting a power budget to the processor/core. This paper initially studies the behavior of different techniques to match a predefined power budget in a CMP processor. While legacy techniques properly work for thread independent/multi-programmed workloads, parallel workloads exhibit the problem of independently adapting the power of each core in a thread dependent scenario. In order to solve this problem we propose a novel mechanism, Power Token Balancing (PTB), aimed at accurately matching an external power constraint by balancing the power consumed among the different cores using a power token-based approach while optimizing the energy efficiency. We can use power (seen as tokens or coupons) from non-critical threads for the benefit of critical threads. PTB runs transparent for thread independent / multiprogrammed workloads and can be also used as a spin lock detector based on power patterns. Results show that PTB matches more accurately a predefined power budget (total energy consumed over the budget is reduced to 8\% for a 16-core CMP) than DVFS with only a 3\% energy increase. Finally, we can trade accuracy on matching the power budget for energy-efficiency reducing the energy a 4% with a 20% of accuracy. Juan M. Cebrian, Juan L. Aragón, Stefanos Kaxiras |
IPDPS | 1 |
| 2011 | Leakage-efficient design of value predictors through state and non-state preserving techniques
Juan M. Cebrian, Juan L. Aragón, José M. García 0001, Stefanos Kaxiras |
J. Supercomput. | 1 |
| 2009 | Efficient microarchitecture policies for accurately adapting to power constraintsabstractIn the past years dynamic voltage and frequency scaling (DVFS) has been an effective technique that allowed microprocessors to match a predefined power budget. However, as process technology shrinks, DVFS becomes less effective (because of the increasing leakage power) and it is getting closer to a point where DVFS won't be useful at all (when static power exceeds dynamic power). In this paper we propose the use of microarchitectural techniques to accurately match a power constraint while maximizing the energy efficiency of the processor. We predict the processor power consumption at a basic block level, using the consumed power translated into tokens to select between different power-saving micro-architectural techniques. These techniques are orthogonal to DVFS so they can be simultaneously applied. We propose a two-level approach where DVFS acts as a coarse-grained technique to lower the average power while microarchitectural techniques remove all the power spikes efficiently. Experimental results show that the use of power-saving microarchitectural techniques in conjunction with DVFS is up to six times more precise, in terms of total energy consumed (area) over the power budget, than using DVFS alone for matching a predefined power budget. Furthermore, in a near future DVFS will become DFS because lowering the supply voltage will be too expensive in terms of leakage power. At that point, the use of power-saving microarchitectural techniques will become even more energy efficient. Juan M. Cebrian, Juan L. Aragón, José M. García 0001, Pavlos Petoumenos, Stefanos Kaxiras |
IPDPS | 1 |
| 2007 | Leakage Energy Reduction in Value Predictors through Static DecayabstractAs process technology advances toward deep submicron (below 90 nm), static power becomes a new challenge to address for energy-efficient high performance processors, especially for large on-chip array structures such as caches and prediction tables. Value prediction appeared as an effective way of increasing processor performance by overcoming data dependences, but at the risk of becoming a thermal hot spot due to the additional power dissipation. This paper proposes the design of low-leakage value predictors by applying static decay techniques in order to disable unused entries from the prediction tables. We explore decay strategies suited for the three most common Value predictors (STP, FCM and DFCM) studying the particular tradeoffs for these prediction structures. Our mechanism reduces VP leakage energy efficiently without compromising VP accuracy nor processor performance. Results show average leakage energy reductions of 52%, 65% and 75% for the STP, DFCM and FCM value predictors, respectively. Juan M. Cebrian, Juan L. Aragón, José M. García 0001 |
IPDPS | 1 |