VLDB 2026 Research / reviewers in the wild / expert
Manuel E. Acacio
dblp:15/5905 · also Manuel E. Acacio Sanchez
· DBLP profile ↗
90ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-0935-4078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 80 · 6 first-author · 14 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUsabstractThe growing complexity and parallelism demands of modern GPU workloads have driven architectural innovations toward asynchronous tile transfers (ATTs) to overlap computation and data movement. While ATT units such as the NVIDIA's Tensor Memory Accelerator (TMA) introduce high-throughput memory transfers, programmers must deal with wavefront specialization, select tile sizes, queue slots, and synchronization primitives, all of which are hardware-specific and workloaddependent. Existing GPU libraries fall short-offering limited ATT support and configurability-so developers still resort to manual exploration of this vast parameter space, which is laborious, error-prone, and fundamentally limits performance portability across GPUs. In this work, we present QuCo (Queue Configurator), a single lightweight hardware unit embedded in the GPU that fully automates the ATT configuration process. Inspired by Blackwell GPU design, QuCo includes a compact RISC-V processor, small memory structures for instructions and data, and a GPU Specification Table (GST) storing key architectural parameters. Using the GST and workload characteristics, along with built-in heuristics, QuCo computes optimal queue configurations at kernel launch. This relieves the programmer of the tedious, time-consuming task of tuning and offline profiling, while simultaneously increasing post-compilation performance portability. Nicolás Meseguer, Daoxuan Xu, Yifan Sun 0002, Michael Pellauer, José L. Abellán, Manuel E. Acacio |
HPCA | 6 |
| 2025 | No Rush in Executing Atomic InstructionsabstractHardware atomic instructions are the building blocks of the synchronization algorithms. Historically, to guarantee atomicity and consistency, they were implemented using memory fences, committing older memory instructions, and draining the store buffer before initiating the execution of atomics. Unfortunately, the use of such memory fences entails huge performance penalties as it implies execution serialization, thus impeding instruction- and memory-level parallelism. The situation, however, seems to have changed recently. Through experiments on $x 86$ machines, we discovered that current $x 86$ processors manage to comply with the x86-TSO requirements while avoiding the performance overhead introduced by fences (fence-free or unfenced implementation). This paves the way to new potential optimizations to atomic instruction execution. In particular, our simulation experiments modeling unfenced atomics reveal that executing atomic instructions as soon as their operands are ready does not always lead to optimal performance. In fact, this increases the time that other threads should wait to obtain the cacheline. In contended scenarios, delaying the execution of the atomic instruction to minimize the time the cacheline is locked provides superior performance. Based on this observation, we present Rush or Wait (RoW), a hardware mechanism to decide when to execute an atomic instruction. The mechanism is based on a contention predictor that estimates if an atomic will access a contended cacheline. Non-contended atomics execute once their operands are ready. Contended atomics, on the contrary, wait to become the oldest memory instruction and to drain the store buffer to execute, minimizing the contention on the accessed cacheline. Our experimental evaluation shows that RoW reduces execution time on average by 9.2% (and up to 43%) compared to a baseline that executes atomics as soon as the operands are ready, and yet it requires a small area overhead (64 bytes). Ashkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras, Alberto Ros 0001 |
HPCA | 3 |
| 2025 | QuFi: Adaptive Tiled Gustavson Output Reuse for Edge Sparse DNN AcceleratorsabstractIn recent years, a myriad of Deep Neural Network (DNN) accelerator architectures have been proposed targeting efficient Sparse Matrix-Sparse Matrix Multiplication (SpMSpM) for the edge. Most focus on the dataflow, i.e., the order in which processing elements perform multiply-accumulate operations, while overlooking the influence of memory structures. For Gustavson-based dataflows, which have been proven to be the most efficient for the mid-sparsity scenarios exhibited by sparse DNN workloads, we find that the memory structure's organization storing partial results significantly affects performance. Different DNN models, even layers within the same model, impose diverse requirements on this structure. Rigid designs, commonly assumed so far, often cause frequent merging operations that degrade performance. In this work, we advocate for the necessity of providing adaptability at this memory structure and propose QuFi, a configurable merging memory structure designed for easy integration into any Gustavson-based accelerator and tailored to edge devices. Implemented as a collection of queues across multiple memory banks, QuFi supports reconfigurability through queue fusion and dynamic clustering while providing high bandwidth. Through a detailed evaluation that considers its inclusion into a state-of-the-art accelerator design for the edge, we show that QuFi provides an average speedup of$1.64 \times$, up to 73 % reduction in off-chip memory traffic, and leads to total accelerator area savings of 17 %. Adrián Navarro, José Cano 0001, José L. Abellán, Manuel E. Acacio |
ICCD | 4 |
| 2025 | Precise characterization of coherence activity in multicores using gem5
Joaquín Ferrer, Juan M. Cebrian, Ricardo Fernández-Pascual, Manuel E. Acacio |
J. Supercomput. | 4 |
| 2024 | Chaining Transactions for Effective Concurrency Management in Hardware Transactional MemoryabstractHardware Transactional Memory (HTM) offers the opportunity to ease parallel programming. However, driven by hardware limitations, commercial implementations eschew the complexity involved in early sophisticated proposals from academia, and, among other things, opt for simple conflict resolution policies that inevitably increase transaction aborts. To increase thread level parallelism, previous works propose conflict resolution schemes that, instead of aborting, add a second level of speculation consisting in using not-yet-committed data from another transaction. This policy, which we refer to as requester-speculates, has not yet been considered in the context of the kind of best-effort HTM support provided by commercial processors. This work proposes CHAining TransactionS (CHATS), a simple yet effective realization of the requester-speculates con-flict resolution policy in which cyclic dependencies between transactions are avoided and the commit ordering respects the dependencies that transactions make once speculative values are communicated. The ultimate result is a best-effort HTM implementation that forces a partial order between transactions in a way that ensures effective utilization of forwarded data and that gets away from the complexity of previous proposals. Simulations using gem5 demonstrate the effectiveness of CHATS in both commercial-like setups and academic state-of-the-art best-effort systems (22% and 16% reduction in execution time, on average, respectively). These improvements are achieved by requiring less than 280 bytes of extra storage. Víctor Nicolás-Conesa, J. Rubén Titos Gil, Ricardo Fernández-Pascual, Manuel E. Acacio, Alberto Ros 0001 |
MICRO | 4 |
| 2023 | CELLO: Compiler-Assisted Efficient Load-Load Ordering in Data-Race-Free RegionsabstractEfficient Total Store Order (TSO) implementations allow loads to execute speculatively out-of-order. To detect order violations, the load queue (LQ) holds all the in-flight loads and is searched on every invalidation and cache eviction. Moreover, in a simultaneous multithreading processor (SMT), stores also search the LQ when writing to cache. LQ searches entail considerable energy consumption. Furthermore, the processor stalls upon encountering the LQ full or when its ports are busy. Hence, the LQ is a critical structure in terms of both energy and performance. In this work, we observe that the use of the LQ could be dramatically optimized under the guarantees of the datarace-free (DRF) property imposed by modern programming languages. To leverage this observation, we propose CELLO, a software-hardware co-design in which the compiler detects memory operations in DRF regions and the hardware optimizes their execution by safely skipping LQ searches without violating the TSO consistency model. Furthermore, CELLO allows removing DRF loads from the LQ earlier, as they do not need to be searched to detect consistency violations. With minimal hardware overhead, we show that an 8-core 2-way SMT processor with CELLO avoids almost all conservative searches to the LQ and significantly reduces its occupancy. CELLO allows i) to reduce the LQ energy expenditure by 33% on average (up to 53%) while performing 2.8% better on average (up to 18.6%) than the baseline system, and ii) to shrink the LQ size from 192 to only 80 entries, reducing the LQ energy expenditure as much as 69% while performing on par with a mainstream LQ implementation. Sawan Singh, Josué Feliu, Manuel E. Acacio, Alexandra Jimborean, Alberto Ros 0001 |
PACT | 3 |
| 2023 | Flexagon: A Multi-dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN ProcessingabstractSparsity is a growing trend in modern DNN models. Francisco Muñoz-Martínez, Raveesh Garg, Michael Pellauer, José L. Abellán, Manuel E. Acacio, Tushar Krishna |
ASPLOS (3) | 5 |
| 2023 | STIFT: A Spatio-Temporal Integrated Folding Tree for Efficient Reductions in Flexible DNN AcceleratorsabstractIncreasing deployment of Deep Neural Networks (DNNs) recently fueled interest in the development of specific accelerator architectures capable of meeting their stringent performance and energy consumption requirements. DNN accelerators can be organized around three separate NoCs, namely distribution, multiplier, and reduction networks (or DN, MN, and RN, respectively) between the global buffer(s) and the compute units (multipliers/adders). Among them, the RN, used to generate and reduce the partial sums produced during DNN processing, is a first-order driver of the area and energy efficiency of the accelerator. RNs can be orchestrated to exploit a Temporal, Spatial or Spatio-Temporal reduction dataflow. Among these, Spatio-Temporal reduction is the one that has shown superior performance. However, as we demonstrate in this work, a state-of-the-art implementation of the Spatio-Temporal reduction dataflow, based on the addition of Accumulators (Ac) to the RN (i.e., RN+Ac strategy), can result into significant area and energy expenses. To cope with this important issue, we propose STIFT (that stands forSpatio-Temporal Integrated Folding Tree) that implements the Spatio-Temporal reduction dataflow entirely on the RN hardware substrate (i.e., without the need for the extra accumulators). STIFT results into significant area and power savings regarding the more complex RN+Ac strategy, at the same time its performance advantage is preserved. Francisco Muñoz-Martínez, José L. Abellán, Manuel E. Acacio, Tushar Krishna |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2023 | Speculative inter-thread store-to-load forwarding in SMT architecturesabstractApplications running on out-of-order cores have benefited for decades of store-to-load forwarding which accelerates communication of store values to loads of the same thread. Despite threads running on a simultaneous multithreading (SMT) core could also access the load queues (LQ) and store queues (SQ) / store buffers (SB) of other threads to allow inter-thread store-to-load forwarding, we have skipped exploiting it because if we allow communication of different SMT threads via their LQs and SQs/SBs, write atomicity may be violated with respect to the outside world beyond the acceptable model of read-own-write-early multiple-copy atomicity (rMCA). In our prior work, we leveraged this idea to propose inter-thread store-to-load forwarding (ITSLF). ITLSF accelerates synchronization and communication of threads running in a simultaneous multi-threading processor by allowing stores in the store-queue of a thread to forward data to loads of another thread running in the same core without violating rMCA. In this work, we extend the original ITSLF mechanism to allow inter-thread forwarding from speculative stores (Spec-ITSLF). Spec-ITSLF allows forwarding store values to other threads earlier, which further accelerates synchronization. Spec-ITSLF outperforms a baseline SMT core by 15%, which is 2% better on average (and up to 5% for the TATP workload) than the original ITSLF mechanism. More importantly, Spec-ITSLF is on par with the original ITSLF mechanism regarding storage overhead but does not need to keep track of the speculative state of stores, which was an important source of overhead and complexity in the original mechanism. Josué Feliu, Alberto Ros 0001, Manuel E. Acacio, Stefanos Kaxiras |
J. Parallel Distributed Comput. | 3 |
| 2022 | Understanding the Design-Space of Sparse/Dense Multiphase GNN dataflows on Spatial AcceleratorsabstractGraph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applications in cloud and HPC. Owing to their unique compute and memory characteristics that come from an interplay between dense and sparse phases of computations, the emergence of recon-figurable dataflow (aka spatial) accelerators offers promise for acceleration by mapping optimized dataflows (i.e., computation order and parallelism) for both phases. The goal of this work is to characterize and understand the design-space of dataflow choices for running GNNs on spatial accelerators in order for mappers or design-space exploration tools to optimize the dataflow based on the workload. Specifically, we propose a taxonomy to describe all possible choices for mapping the dense and sparse phases of GNN inference, spatially and temporally over a spatial accelerator, capturing both the intra-phase dataflow and the inter-phase (pipelined) dataflow. Using this taxonomy, we do deep-dives into the cost and benefits of several dataflows and perform case studies on implications of hardware parameters for dataflows and value of flexibility to support pipelined execution. Raveesh Garg, Eric Qin 0001, Francisco Muñoz-Martínez, Robert Guirado, Akshay Jain 0001, Sergi Abadal, José L. Abellán, Manuel E. Acacio, Eduard Alarcón, Sivasankaran Rajamanickam, Tushar Krishna |
IPDPS | 8 |
| 2022 | Analysis of the Interactions Between ILP and TLP With Hardware Transactional MemoryabstractHardware Transactional Memory (HTM) allows the use of transactions by programmers, making parallel programming easier and theoretically obtaining the performance of fine-grained locks. However, transactions can abort for a variety of reasons, resulting in the squash of speculatively executed instructions and the consequent loss in both performance and energy efficiency. Among the different sources of abort, conflicting concurrent accesses to the same shared memory locations from different transactions are often the prevalent cause.In this work, we characterize, for the first time to the best of our knowledge, how the aggressiveness of the cores in terms of exploiting instruction-level parallelism can interact with thread-level speculation support brought by HTM systems. We observe that altering the size of the structures that support out-of-order and speculative execution changes the number of aborts produced in the execution of transactional workloads on a best-effort HTM implementation. Our results show that a small number of powerful cores is more suitable for high-contention scenarios, whereas under low contention it is preferable to use a larger number of less aggressive cores. In addition, an aggressive core can lead to performance loss in medium-contention scenarios due to an increase in the number of aborts. We conclude that depending on contention, a careful choice over processor aggressiveness can reduce abort ratios. Víctor Nicolás-Conesa, J. Rubén Titos Gil, Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
PDP | 5 |
| 2022 | Analysing software prefetching opportunities in hardware transactional memory
Marina Shimchenko, J. Rubén Titos Gil, Ricardo Fernández-Pascual, Manuel E. Acacio, Stefanos Kaxiras, Alberto Ros 0001, Alexandra Jimborean |
J. Supercomput. | 4 |
| 2022 | DeTraS: Delaying Stores for Friendly-Fire Mitigation in Hardware Transactional MemoryabstractCommercial Hardware Transactional Memory (HTM) systems are best-effort designs that leverage the coherence substrate to detect conflicts eagerly. Resolving conflicts in favor of the requesting core is the simplest option for ensuring deadlock freedom, yet it is prone to livelocks. In this work, we propose and evaluate DeTraS (Delayed Transactional Stores), an HTM-aware store buffer design aimed at mitigating such livelocks. DeTraS takes advantage of the fact that modern commercial processors implement a large store buffer, and uses it to prevent transactional stores predicted to conflict from performing early in the transaction. By leveraging existing processor structures, we propose a simple design that improves the ability of requester-wins HTM systems to achieve forward progress in spite of high contention while side-stepping the performance penalty of falling back to mutual exclusion. With just over 50 extra bytes, DeTraS captures the advantages of lazy conflict management without the complexity brought into the coherence fabric by commit arbitration schemes nor the relaxation of the single-writer invariant of prior works. Through detailed simulations of a 16-core tiled CMP using gem5, we demonstrate that DeTraS brings reductions in average execution time of 25 percent when compared to an Intel RTM-like design. J. Rubén Titos Gil, Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | ITSLF: Inter-Thread Store-to-Load Forwardingin Simultaneous MultithreadingabstractIn this paper, we argue that, for a class of fine-grain, synchronization-intensive, parallel workloads, it is advantageous to consolidate synchronization and communication as much as possible among the threads of simultaneous multithreading (SMT) cores. While, today, the shared L1 is the closest coherent level where synchronization and communication between SMT threads can take place, we observe that there is an even closer shared level, entirely inside a single core. This level comprises the load queues (LQ) and store queues (SQ) / store buffers (SB) of the SMT threads and to the best of our knowledge it has never been used as such. The reason is that if we allow communication of different SMT threads via their LQs and SQs/SBs, i.e., inter-thread store-to-load forwarding (ITSLF), we violate write atomicity with respect to the outside world, beyond the acceptable model of read-own-write-early multiple-copy atomicity (rMCA). Josué Feliu, Alberto Ros 0001, Manuel E. Acacio, Stefanos Kaxiras |
MICRO | 3 |
| 2021 | A novel network fabric for efficient spatio-temporal reduction in flexible DNN acceleratorsabstractIncreasing deployment of Deep Neural Networks (DNNs) in a myriad of applications, has recently fueled interest in the development of specific accelerator architectures capable of meeting their stringent performance and energy consumption requirements. Francisco Muñoz-Martínez, José L. Abellán, Manuel E. Acacio, Tushar Krishna |
NOCS | 3 |
| 2020 | PfTouch: Concurrent page-fault handling for Intel restricted transactional memory
J. Rubén Titos Gil, Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
J. Parallel Distributed Comput. | 4 |
| 2020 | Concurrent Irrevocability in Best-Effort Hardware Transactional MemoryabstractExisting best-effort requester-wins implementations of transactional memory must resort to non-speculative execution to provide forward progress in the presence of transactions that exceed hardware capacity, experience page faults or suffer high-contention leading to livelocks. Current approaches to irrevocability employ lock-based synchronization to achieve mutual exclusion when executing a transaction non-speculatively, conservatively precluding concurrency with any other transactions in order to guarantee atomicity at the cost of degrading performance. In this article, we propose a new form of concurrent irrevocability whose goal is to minimize the loss of concurrency paid when transactions resort to irrevocability to complete. By enabling optimistic concurrency control also during non-speculative execution of a transaction, our proposal allows for higher parallelism than existing schemes. We describe the extensions to the instruction set to provide concurrent irrevocable transactions as well as the architectural extensions required to realize them on a best-effort HTM system without requiring any modification to the cache coherence protocol. Our evaluation shows that our proposal achieves an average reduction of 12.5 percent in execution time across the STAMP benchmarks, with 15.8 percent on average for highly contended workloads. J. Rubén Titos Gil, Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Foreword to the Special Issue on Processors, Interconnects, Storage, and Caches for Exascale SystemsabstractExascale computing constitutes nowadays a significant challenge both for the academia and the industry. Although traditional computer systems continue to make important advances, achieving exascale computing requires mass customization. With this aim, several ongoing research projects are focusing on different architectural (computing boards or nodes, interconnects, storage, etc) issues of future exascale systems. Most of them devise heterogeneous computing boards consisting of CPUs (high performance and/or low power), FPGAs, GPUs, etc, sharing a common memory hierarchy. In this context, efficient intra- and inter-board within the same rack and inter-rack interconnect with the memory hierarchies are required. Also, performance and reliability design constraints for exascale storage systems rise significant challenges for HPC system designers. High performance I/O must be also faced because storing and retrieving such large amounts of data can greatly affect the overall performance of applications. Finally, it is important to characterize the demands that exascale applications exert on the different components of an exascale system. The goal of this special issue is to promote research on all of these aspects related to exascale computing. Six papers that address several components of exascale systems were carefully selected from open submissions. The six manuscripts included in this special issue cover different aspects of an exascale system. Lant et al1 present a network interface architecture and networking infrastructure, designed to sit inside the FPGA fabric of a cutting-edge heterogeneous MPSoC (Multi-Processor System-on-Chip), enabling networks of these devices to communicate within both a distributed and shared memory context, with reduced need for costly software networking system calls. This work presents in detail the factors that influenced the implementation and system prototype-based upon the use of Xilinx Zynq Ultrascale+ and discusses the main design decisions and implementation challenges. Crespo et al2 emphasize the need of interconnect technologies alternative to the classical electrical one. This work focuses on silicon photonics and highlights practical challenges that must be met to enable the adoption of this technology in building efficient, extreme-scale interconnection networks. In particular, they show that signal loss sources, suffered mainly due to waveguide crossings and propagation, play a critical role in photonic exascale network designs as they constrain the ability to perform data transmission in an effective as network size increases. Also focused on the interconnection network of an exascale system, Duro et al3 conduct an extensive simulation study using realistic photonic network configurations with synthetic and realistic traffic and show that, compared to electrical networks, optical networks can reduce the execution time of the studied real workloads in almost one order of magnitude. The study is performed from an architectural perspective, and the authors state that the photonic configuration highly impacts on the network performance, being the bandwidth per channel and the message length the most important parameters. Piernas and González-Férez4 address the scalability of file systems aimed to exascale systems. In particular, they describe how they have implemented the support for data objects in their previously proposed Fusion Parallel File System (FPFS). They show that the utilization of a unified data and metadata server (an enhanced object-based storage device or OSD+) provides FPFS with a competitive advantage over other file systems like Lustre or OrangeF, which brings higher performance in some file operations. Metadata-intensive workloads are used to stress the network traffic and analyze the scalability of the file systems. Pascual et al5 investigate alternatives for the storage subsystem of a novel exascale-capable system with special emphasis on how allocation strategies would affect the overall performance. They consider several aspects of data-aware allocation (such as the effect of spatial and temporal locality, the affinity of data to storage sources, and the network-level traffic prioritization for different types of flows) and show that scheduling policies exposing data-locality information can be essential for the appropriate utilization of future large-scale systems. They also found that the distributed storage system they implement can outperform traditional SAN architectures, even with a much smaller (in terms of I/O servers) back-end. Finally, Castro et al6 present an energy study on critical parameters for the deployment of CNNs on flagship image and video applications, ie, object recognition and people identification by gait, respectively. Their experimental results on a multi-GPU server endowed with twin Maxwell and twin Pascal Titan X GPUs demonstrate that energy correlates with performance and that Pascal may have up to 40% gains versus Maxwell. Larger batch sizes extend performance gains and energy savings but accuracy must be watched, which sometimes shows a preference for small batches. The manuscripts presented in this special issue provide insights into several cutting-edge aspects of exascale computing. We believe that the main contributions presented in these manuscripts are timely and important. We hope that readers can benefit from these research manuscripts and contribute to these rapidly growing areas. Manuel E. Acacio is a Full Professor of computer architecture and technology at the University of Murcia, Spain. He obtained his PhD degree in Computer Science in March 2003. Before, in the summer of 2002, he worked as a summer intern at IBM TJ Watson, Yorktown Heights (NY). Currently, Prof. Acacio leads the Computer Architecture & Parallel Systems (CAPS) research group at the University of Murcia. He is author of more than 100 papers in refereed international conferences and journals. As well, he has served as a committee member of important conferences, ICPP and IPDPS among others. His research interests are focused on the architecture of multiprocessor systems. From April 2011 to April 2015, Prof. Acacio served as an associate editor of IEEE Transactions on Parallel and Distributed Systems International Journal; since August 2016, he is a member of the editorial board of MPDI Computers Int'l Journal; and more recently, since September 2018, he serves as an academic editor in the editorial board of Hindawi Scientific Programming journal. He is also a member of the board of distinguished reviewers of ACM Transactions on Architecture and Code Optimization Int'l Journal since May 2014. Julio Sahuquillo is a Full Professor with the Department of Computer Engineering at the Universitat Politècnica de València. He has enjoyed a postdoctoral research stay with Prof Antonio Gónzalez, former director of Intel Barcelona. He has taught several courses on computer organization and architecture. He has co-authored more than 150 refereed conference and journal papers. His current research interests include multi- and manycore processors, memory hierarchy design, cache coherence, GPU architecture, resource management in the cloud, and architecture-aware scheduling. In these topics, he has advised more than 10 PhD Theses and has been the principal investigator of competitive Spanish domestic projects, international European projects, and projects with international companies. He has participated in the organization of about 30 conferences (HPCC, Euro-Par, HPCS, etc.) in different positions: Publicity Chair, Local Organizing Committee, Workshop Co-Chair, and Special Session Co-Chair. He participates assiduously in the PC of major Computer Architecture conferences. He is a member of the IEEE and the IEEE Computer Society. The guess editors would like to thank all the authors who made valuable contributions to this special issue. We also thank the reviewers for their detailed review reports that have helped to further enhance the manuscripts originally submitted. Finally, we would like to express our sincere gratitude to Prof Geoffrey Fox, the editor-in-chief, for having provided us with the opportunity to edit this special issue in the international journal of Concurrency and Computation: Practice and Experience, as well as for his assistance throughout all the review process. Manuel E. Acacio, Julio Sahuquillo |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | InsideNet: A tool for characterizing convolutional neural networks
Francisco Muñoz-Martínez, José L. Abellán, Manuel E. Acacio |
Future Gener. Comput. Syst. | 3 |
| 2019 | Way Combination for an Adaptive and Scalable Coherence DirectoryabstractThis manuscript opens the way to a new class of coherence directory structures that are based on the brand-new concept of way combining. A Way-Combining Directory (WC-dir) builds on a typical sparse directory but allows to take advantage of several ways in the same set to codify the sharing information of each memory block. The result is a sparse directory with variable effective associativity per set and variable length entries, thus being able to dynamically adapt the directory structure to the particular requirements of each application. In particular, our proposal uses just enough bits per entry to store a single pointer, which is optimal for the common case of having just one sharer. For those addresses that have more than one sharer, we have observed that in the majority of cases extra bits could be taken from other empty ways in the same set. All in all, our proposal minimizes the storage overheads without losing the flexibility to adapt to several sharing degrees and without the complexities of other previously proposed techniques. Detailed simulations of a 128-core multicore architecture running benchmarks from PARSEC-3.0 and SPLASH-3 demonstrate that WC-dir can closely approach the performance of a non-scalable bit vector sparse directory, beating the state-of-the-art Scalable Coherence Directory (SCD) and Pool directory proposals. J. Rubén Titos Gil, Antonio Flores, Ricardo Fernández-Pascual, Alberto Ros 0001, Salvador Petit, Julio Sahuquillo, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2018 | SAWS: Simple and Adaptive Warp Scheduling for Improved Performance in Throughput ProcessorsabstractIn this work, we address the challenge of designing an efficient warp scheduler for throughput processors by proposing SAWS (Simple and Adaptive Warp Scheduler). Differently from previous approaches which target a particular type of applications, SAWS considers several simple scheduling algorithms and tries to use the one that best fits each application or phase within an application. Through detailed simulations we demonstrate that a practical implementation of SAWS can obtain IPC values that closely match the best scheduling algorithm in each case. Francisco Muñoz-Martínez, Manuel E. Acacio |
PDP | 2 |
| 2018 | Photonic-based express coherence notifications for many-core CMPs
José L. Abellán, Eduardo Padierna, Alberto Ros 0001, Manuel E. Acacio |
J. Parallel Distributed Comput. | 4 |
| 2018 | Parallel implementations of the 3D fast wavelet transform on a Raspberry Pi 2 cluster
Gregorio Bernabé, Raúl Hernández, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2017 | Way-combining directory: an adaptive and scalable low-cost coherence directoryabstractToday, general-purpose commercial multicores approaching one hundred cores are already a reality and even thousand core chips are being prototyped. Maintaining coherence across such a high number of cores in these manycore architectures requires careful design of the coherence directory used to keep track of current locations of the memory blocks at the private cache level. In this work we propose a novel organization for the coherence directory that builds on the brand-new concept of way combining. Particularly, our proposal employs just one pointer per entry, which is optimal for the common case of having just one sharer. For those addresses that require more than one pointer, we have observed that in the majority of cases extra pointers could be taken from other empty ways in the same set. Thus, our proposal minimizes the storage overheads without losing the flexibility to adapt to several sharing degrees and without the complexities of other previously proposed techniques. Through detailed simulations of a 128-core architecture, we show that the way-combining directory closely approaches the performance of a non-scalable bit-vector sparse directory, and beats other scalable state-of-the-art proposals. J. Rubén Titos Gil, Antonio Flores, Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
ICS | 5 |
| 2017 | A dedicated private-shared cache design for scalable multiprocessorsabstractSummary Most modern architectures are based on a shared‐memory design. Correctness of these architectures is ensured by means of coherence protocols and consistency models. However, performance and scalability of shared‐memory systems is usually constrained by the amount and size of the messages used to keep the memory subsystem coherent. This is not only important in high performance computing, but also in low power embedded systems, specially if coherence is required between different components of the system‐on‐chip. We argue that using the same mechanism to keep coherence for all memory accesses can be counterproductive, because it incurs unnecessary overhead for data addresses that would remain coherent after the access (i.e., private data and read‐only shared data). This paper proposes the use of dedicated caches for two different kinds of data (i) data that can be accessed without contacting other nodes and (ii) modifiable shared data. The private cache (L1P) will be independent for each core and will store private data and read‐only shared data. On the other hand, the shared cache (L1S), will be logically shared but physically distributed for all cores. With this design, we can significantly simplify the coherence protocol, reduce the on‐chip area requirements and reduce invalidation time. However, this dedicated cache design requires a classification mechanism to detect the nature of the data that is being accessed. Results show two drawbacks to this approach: first, the accuracy of the classification mechanism has a huge impact on performance. Second, a traditional interconnection network is not optimal for accessing the L1S, increasing register‐to‐cache latency when accessing shared data. Copyright © 2016 John Wiley & Sons, Ltd. Juan M. Cebrian, Ricardo Fernández-Pascual, Alexandra Jimborean, Manuel E. Acacio, Alberto Ros 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | To be silent or not: on the impact of evictions of clean data in cache-coherent multicores
Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2016 | Are distributed sharing codes a solution to the scalability problem of coherence directories in manycores? An evaluation study
Ricardo Fernández-Pascual, Alberto Ros 0001, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2015 | Early Experiences with Separate Caches for Private and Shared DataabstractShared-memory architectures have become predominant in modern multi-core microprocessors in all market segments, from embedded to high performance computing. Correctness of these architectures is ensured by means of coherence protocols and consistency models. Performance and scalability of shared-memory systems is usually limited by the amount and size of the messages used to keep the memory subsystem coherent. Moreover, we believe that blindly keeping coherence for all memory accesses can be counterproductive, since it incurs in unnecessary overhead for data that will remain coherent after the access. Having this in mind, in this paper we propose the use of dedicated caches for private (+shared read-only) and shared data. The private cache (L1P) will be independent for each core while the shared cache (L1S) will be logically shared but physically distributed for all cores. This separation should allow us to simplify the coherence protocol, reduce the on-chip area requirements and reduce invalidation time with minimal impact on performance. The dedicated cache design requires a classification mechanism to detect private and shared data. In our evaluation we will use a classification mechanism that operates at the operating system (OS) level (page granularity). Results show two drawbacks to this approach: first, the selected classification mechanism has too many false positives, thus becoming an important limiting factor. Second, a traditional interconnection network is not optimal for accessing the L1S, and a custom network design is needed. These drawbacks lead to important performance degradation due to the additional latency when accessing the shared data. Juan M. Cebrian, Alberto Ros 0001, Ricardo Fernández-Pascual, Manuel E. Acacio |
e-Science | 4 |
| 2015 | Adaptive Selection of Cache Indexing Bits for Removing Conflict MissesabstractThe design of cache memories is a crucial part of the design cycle of a modern processor, since they are able to bridge the performance gap between the processor and the memory. Unfortunately, caches with low degrees of associativity suffer a large amount of conflict misses. Although by increasing their associativity a significant fraction of these misses can be removed, this comes at a high cost in both power, area, and access time. In this work, we address the problem of high number of conflict misses in low-associative caches, by proposing an indexing policy that adaptively selects the bits from the block address used to index the cache. The basic premise of this work is that the non-uniformity in the set usage is caused by a poor selection of the indexing bits. Instead, by selecting at run time those bits that disperse the working set more evenly across the available sets, a large fraction of the conflict misses (85 percent, on average) can be removed. This leads to IPC improvements of 10.9 percent for the SPEC CPU2006 benchmark suite. By having less accesses in the L2 cache, our proposal also reduces the energy consumption of the cache hierarchy by 13.2 percent. These benefits come with a negligible area overhead. Alberto Ros 0001, Polychronis Xekalakis, Marcelo Cintra, Manuel E. Acacio, José M. García 0001 |
IEEE Trans. Computers | 4 |
| 2015 | Fast and efficient commits for Lazy-Lazy hardware transactional memory
Epifanio Gaona-Ramírez, José L. Abellán, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2015 | DASC-DIR: a low-overhead coherence directory for many-core processors
Alberto Ros 0001, Manuel E. Acacio |
J. Supercomput. | 2 |
| 2014 | Selective dynamic serialization for reducing energy consumption in hardware transactional memory systems
Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
J. Supercomput. | 4 |
| 2014 | ZEBRA: Data-Centric Contention Management in Hardware Transactional MemoryabstractTransactional contention management policies show considerable variation in relative performance with changing workload characteristics. Consequently, incorporation of fixed-policy Transactional Memory (TM) in general purpose computing systems is suboptimal by design and renders such systems susceptible to pathologies. Of particular concern are Hardware TM (HTM) systems where traditional designs have hardwired policies in silicon. Adaptive HTMs hold promise, but pose major challenges in terms of design and verification costs. In this paper, we present the ZEBRA HTM design, which lays down a simple yet high-performance approach to implement adaptive contention management in hardware. Prior work in this area has associated contention with transactional code blocks. However, we discover that by associating contention with data (cache blocks) accessed by transactional code rather than the code block itself, we achieve a neat match in granularity with that of the cache coherence protocol. This leads to a design that is very simple and yet able to track closely or exceed the performance of the best performing policy for a given workload. ZEBRA, therefore, brings together the inherent benefits of traditional eager HTMs-parallel commits-and lazy HTMs-good optimistic concurrency without deadlock avoidance mechanisms-, combining them into a low-complexity design. J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Towards Efficient Dynamic LLC Home Bank Mapping with NoC-Level Support
Mario Lodde, José Flich, Manuel E. Acacio |
Euro-Par | 3 |
| 2013 | On the design of energy-efficient hardware transactional memory systemsabstractSUMMARY Transactional memory is currently being advocated as a promising alternative to lock‐based synchronization because it simplifies multithreaded programming. In this way, future many‐core chip multiprocessor architectures may need to provide hardware support for transactional memory. On the other hand, energy consumption constitutes nowadays a first class consideration in multicore processor designs. In this work, we characterize the performance and energy consumption of two well‐known hardware transactional memory systems that employ opposite policies for data versioning and conflict management. More specifically, we compare a LogTM‐SEeager‐eagersystem and a version of the Scalable Transactional Coherence and Consistencylazy‐lazysystem that enable parallel commits. To do so, we extended the Multifacet GEMS simulator to estimate the energy consumed in the on‐chip caches according to CACTI and used the interconnection network energy model given by Orion 2. Results show that the energy consumption of the eager‐eager system is 38% higher in average than in the lazy‐lazy case, whereas performance differences between the two systems are 26% in average. We found that even though lazy‐lazy beats eager‐eager on average, there are considerable deviations in performance depending on the particular characteristics of each application and the settings of both systems. Finally, from this characterization, we observe that a significant part of the energy consumed in some applications in eager‐eager is spent on the back‐off delay phase and explore more energy‐efficient hardware back‐off mechanisms. For lazy‐lazy systems, the way in which memory lines are assigned to the L2 cache banks affects the number of parallel commits in some applications, and we study an alternative fine‐grained assignment. Copyright © 2012 John Wiley & Sons, Ltd. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Design of an efficient communication infrastructure for highly contended locks in many-core CMPs
José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
J. Parallel Distributed Comput. | 3 |
| 2013 | Efficient Eager Management of Conflicts for Scalable Hardware Transactional MemoryabstractThe efficient management of conflicts among concurrent transactions constitutes a key aspect that hardware transactional memory (HTM) systems must achieve. Scalable HTM proposals so far inherit the cache-based style of conflict detection typically found in bus-based systems, largely unaware of the interactions between transactions and directory coherence. In this paper, we demonstrate that the traditional approach of detecting conflicts at the private cache levels is inefficient when used in the context of a directory protocol. We find that the use of the directory as a mere router of coherence requests restricts the throughput of conflict detection, and show how it becomes a bottleneck under high contention. This paper proposes a scheme for conflict detection that decouples conflict detection from cache coherence in order to overcome pathological situations that degrade the performance of an eager HTM system. Our scheme places bookkeeping metadata at the directory, introducing it as a separate hardware module that leaves the coherence protocol unmodified. In comparison to a state-of-the-art eager HTM system, our design handles contention more efficiently, minimizes the performance degradation of false positives for signatures of similar hardware cost, and reduces the network traffic generated. J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | Eager Beats Lazy: Improving Store Management in Eager Hardware Transactional MemoryabstractHardware transactional memory (HTM) designs are very sensitive to the manner in which speculative updates from transactions are handled in the system. This study highlights how the lack of effective techniques for store management results in a quick degradation in the performance of eager HTM systems with increasing contention and, thus, lends credence to the belief that eager designs do not perform as well as their lazy counterparts when conflicts abound. In this work, we present two simple ways to improve handling of speculative stores--a way to effectively manage lines that exhibit migratory sharing and a way to hide store latency, particularly for those stores that target contended cache lines owned by other concurrent transactions. These two mechanisms yield substantial improvements in execution time when running applications with high contention, allowing eager designs to exceed the performance of lazy ones. Interestingly, the benefits that accrue from these enhancements can be at par with those achieved using more complex system-wide HTM techniques. Coupled with the fact that eager designs are easier to integrate into cache coherent architectures than lazy ones, we claim that with judicious management of stores they represent a more compelling design alternative. J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Design of a collective communication infrastructure for barrier synchronization in cluster-based nanoscale MPSoCsabstractBarrier synchronization is a key programming primitive for shared memory embedded MPSoCs. As the core count increases, software implementations cannot provide the needed performance and scalability, thus making hardware acceleration critical. In this paper we describe an interconnect extension implemented with standard cells and with a mainstream industrial toolflow. We show that the area overhead is marginal with respect to the performance improvements of the resulting hardware-accelerated barriers. We integrate our HW barrier into the OpenMP programming model and discuss synchronization efficiency compared with traditional software implementations. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, Davide Bertozzi, Daniele Bortolotti, Andrea Marongiu, Luca Benini |
DATE | 3 |
| 2012 | Dynamic Last-Level Cache Allocation to Reduce Area and Power Overhead in Directory Coherence Protocols
Mario Lodde, José Flich, Manuel E. Acacio |
Euro-Par | 3 |
| 2012 | π-TM: Pessimistic invalidation for scalable lazy hardware transactional memoryabstractLazy hardware transactional memory has been shown to be more efficient at extracting available concurrency than its eager counterpart. However, it poses scalability challenges at commit time as existence of conflicts among concurrent transactions is not known prior to commit. Non-conflicting transactions may have to wait before committing, severely affecting performance in certain workloads. Early conflict detection can be employed to allow such transactions to commit simultaneously. In this paper we show that the potential of this technique has not yet been fully utilized, with design choices in prior work severely burdening common-case transactional execution to avoid some relatively uncommon correctness concerns. The paper quantifies the severity of the problem and develops μ-TM, an early conflict detection - lazy conflict resolution design. This design highlights how, with modest extensions to existing directory-based coherence protocols, information regarding possible conflicts can be effectively used to achieve true parallelism at commit without burdening the common-case. We leverage the observation that contention is typically seen on only a small fraction of shared data accessed by coarse-grained transactions. Pessimistic invalidation of such lines when committing or aborting, therefore, enables fast common-case execution. Our results show that μ-TM performs consistently well and, in particular, far better than previous work on early conflict detection in lazy HTM. We also identify a pathological scenario that lazy designs with early conflict detection suffer from and propose a simple hardware workaround to sidestep it. Anurag Negi, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Per Stenström |
HPCA | 3 |
| 2012 | ASCIB: adaptive selection of cache indexing bits for removing conflict missesabstractThe design of cache memories is a crucial part of the design cycle of a modern processor. Unfortunately, caches with low degrees of associativity suffer a large amount of conflict misses, while high-associative caches consume more power per access. We propose ASCIB, a simple technique able to dynamically adjust the bits used for cache indexing so as to minimize conflict misses. By selecting at run time the bits that disperse the working set more evenly across the available sets, ASCIB removes 73% of the conflict misses on average. This results in an improvement in energy efficiency by 17% on average. Alberto Ros 0001, Polychronis Xekalakis, Marcelo Cintra, Manuel E. Acacio, José M. García 0001 |
ISLPED | 4 |
| 2012 | Heterogeneous NoC Design for Efficient Broadcast-based Coherence Protocol SupportabstractChip Multiprocessor Systems (CMPs) rely on a cache coherency protocol to maintain memory access coherence between cached data and main memory. The Hammer coherency protocol is appealing as it eliminates most of the space overhead when compared to a directory protocol. However, it generates much more traffic, thus stressing the NoC and having worse performance in terms of power consumption. When using a NoC with built-in broadcast support network utilization is lowered but does not solve completely the problem as acknowledgment messages are still sent from each core to the memory access requestor. In this paper we propose a simple control network that collects the acknowledgement messages and delivers them with a bounded and fixed latency, thus relieving the NoC from a large amount of messages. Experimental results demonstrate on a 16-tile system with the control network that execution time improves up to 17%, with an average improvement of about 7.5%. The control network has negligible impact on area when compared to the switches. Mario Lodde, José Flich, Manuel E. Acacio |
NOCS | 3 |
| 2012 | Dynamic Serialization: Improving Energy Consumption in Eager-Eager Hardware Transactional Memory SystemsabstractIn the search for new paradigms to simplify multithreaded programming, Transactional Memory (TM) is currently being advocated as a promising alternative to deadlock-prone lock-based synchronization. In this way, future many-core CMP architectures may need to provide hardware support for TM. On the other hand, power dissipation constitutes a first class consideration in multicore processor designs. In this work, we propose Dynamic Serialization (DS) as a new technique to improve energy consumption without degrading performance in applications with conflicting transactions. Our proposal, which is implemented on top of a hardware transactional memory system with an eager conflict management policy, detects and serializes conflicting transactions dynamically. Particularly, in case of conflict one transaction is allowed to continue whilst the rest are completely stalled. Once the executing transaction has finished it wakes up several of the stalling transactions. This brings important benefits in terms of energy consumption due to the reduction in the amount of wasted work that DS implies. Results for a 16-core CMP show that Dynamic Serialization obtains reductions of 10% on average in energy consumption (more than 20% in high contention scenarios) without affecting, on average, execution time. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Manuel E. Acacio, Juan Fernández Peinador |
PDP | 3 |
| 2012 | Using Heterogeneous Networks to Improve Energy Efficiency in Direct Coherence Protocols for Many-Core CMPsabstractDirect coherence protocols have been recently proposed as an alternative to directory-based protocols to keep cache coherence in many-core CMPs. Differently from directory-based protocols, in direct coherence the responsible for providing the requested data in case of a cache miss (i.e., the owner cache) is also tasked with keeping the updated directory information and serializing the different accesses to the block by all cores. This way, these protocols send requests directly to the owner cache, thus avoiding the indirection caused by accessing a separate directory (usually in the home node). A hints mechanism ensures a high hit rate when predicting the current owner of a block for sending requests, but at the price of significantly increasing network traffic, and consequently, energy consumption. In this work, we show how using a heterogeneous interconnection network composed of two kinds of links is enough to drastically reduce the energy consumed by hint messages, obtaining significant improvements in energy efficiency. Alberto Ros 0001, Ricardo Fernández-Pascual, Manuel E. Acacio |
SBAC-PAD | 3 |
| 2012 | Hardware transactional memory with software-defined conflictsabstractIn this paper we investigate the benefits of turning the concept of transactional conflict from its traditionally fixed definition into a variable one that can be dynamically controlled in software. We propose the extension of the atomic language construct with an attribute that specifies the definition of conflict, so that programmers can write code which adjusts what kinds of conflicts are to be detected, relaxing or tightening the conditions according to the forms of interference that can be tolerated by a particular algorithm. Using this performance-motivated construct, specific conflict information can be associated with portions of code, as each transaction is provided with a local definition that applies while it executes. We find that defining conflicts in software makes possible the removal of dependencies which arise as a result of the coarse synchronization style encouraged by the TM programming model. We illustrate the use of the proposed construct in a variety of use cases with real applications, showing how programmers can take advantage of their knowledge about the problem and other global information not available at run-time. We describe how to implement a hardware TM design that utilizes this software construct. Our experiments reveal that leveraging software-defined conflicts, the programmer is able to achieve significant reductions in the number of aborts--over 50% for most applications. At 16 threads, our system with software-defined conflicts outperforms LogTM-SE in nearly all benchmarks, reaching an average reduction in execution time of 18%. J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Tim Harris 0001, Adrián Cristal, Osman S. Unsal, Ibrahim Hur, Mateo Valero |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Extending Magny-Cours Cache CoherenceabstractOne cost-effective way to meet the increasing demand for larger high-performance shared-memory servers is to build clusters with off-the-shelf processors connected with low-latency point-to-point interconnections like HyperTransport. Unfortunately, HyperTransport addressing limitations prevent building systems with more than eight nodes. While the recent High-Node Count HyperTransport specification overcomes this limitation, recently launched twelve-core Magny-Cours processors have already inherited it and provide only 3 bits to encode the pointers used by the directory cache which they include to increase the scalability of their coherence protocol. In this work, we propose and develop an external device to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the directory cache. Evaluation results for systems with up to 32 nodes show that the performance offered by our solution scales with the number of nodes, enhancing the directory cache effectiveness by filtering additional messages. Particularly, we reduce execution time by 47 percent in a 32-die system with respect to the 8-die Magny-Cours configuration. Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato |
IEEE Trans. Computers | 5 |
| 2012 | Stencil computations on heterogeneous platforms for the Jacobi method: GPUs versus Cell BE
José M. Cecilia, José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, José M. García 0001, Manuel Ujaldon |
J. Supercomput. | 4 |
| 2012 | Efficient Hardware Barrier Synchronization in Many-Core CMPsabstractTraditional software-based barrier implementations for shared memory parallel machines tend to produce hotspots in terms of memory and network contention as the number of processors increases. This could limit their applicability to future many-core CMPs in which possibly several dozens of cores would need to be synchronized efficiently. In this work, we develop GBarrier, a hardware-based barrier mechanism especially aimed at providing efficient barriers in future many-core CMPs. Our proposal deploys a dedicated G-line-based network to allow for fast and efficient signaling of barrier arrival and departure. Since GBarrier does not have any influence on the memory system, we avoid all coherence activity and barrier-related network traffic that traditional approaches introduce and that restrict scalability. Through detailed simulations of a 32-core CMP, we compare GBarrier against one of the most efficient software-based barrier implementations for a set of kernels and scientific applications. Evaluation results show average reductions of 54 and 21 percent in execution time, 53 and 18 percent in network traffic, and also 76 and 31 percent in the energy-delay2product metric for the full CMP when the kernels and scientific applications, respectively, are considered. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | Pi-TM: Pessimistic Invalidation for Scalable Lazy Hardware Transactional MemoryabstractLazy hardware transactional memory (HTM) allows better utilization of available concurrency in transactional workloads than eager HTM, but poses challenges at commit time due to the requirement of en-masse publication of speculative updates to global system state. Early conflict detection can be employed in lazy HTM designs to allow non-conflicting transactions to commit in parallel. Though this has the potential to improve performance, it has not been utilized effectively so far. Prior work in the area burdens common-case transactional execution severely to avoid some relatively uncommon correctness concerns. In this work we investigate this problem and introduce a novel design, π-TM, which eliminates this problem. π-TM uses modest extensions to existing directory-based cache coherence protocols to keep a record of conflicting cache lines as a transaction executes. This information allows a consistent cache state to be maintained when transactions commit or abort. We observe that contention is typically seen only on a small fraction of shared data accessed by coarse-grained transactions. In π-TM early conflict detection mechanisms imply additional work only when such contention actually exists. Thus, the design is able to avoid expensive core-to-core and core-to-directory communication for a large part of transactionally accessed data. Our evaluation shows major performance gains when compared to other HTM designs in this class and competitive performance when compared to more complex lazy commit schemes. Anurag Negi, Per Stenström, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001 |
PACT | 4 |
| 2011 | Eager Meets Lazy: The Impact of Write-Buffering on Hardware Transactional MemoryabstractHardware transactional memory (HTM) systems have been studied extensively along the dimensions of speculative versioning and contention management policies. The relative performance of several designs policies has been discussed at length in prior work within the framework of scalable chip-multiprocessing systems. Yet, the impact of simple structural optimizations like write-buffering has not been investigated and performance deviations due to the presence or absence of these optimizations remains unclear. This lack of insight into the effective use and impact of these interfacial structures between the processor core and the coherent memory hierarchy forms the crux of the problem we study in this paper. Through detailed modeling of various write-buffering configurations we show that they play a major role in determining the overall performance of a practical HTM system. Our study of both eager and lazy conflict resolution mechanisms in a scalable parallel architecture notes a remarkable convergence of the performance of these two diametrically opposite design points when write buffers are introduced and used well to support the common case. Mitigation of redundant actions, fewer invalidations on abort, latency-hiding and prefetch effects contribute towards reducing execution times for transactions. Shorter transaction durations also imply a lower contention probability, thereby amplifying gains even further. The insights, related to the interplay between buffering mechanisms, system policies and workload characteristics, contained in this paper clearly distinguish gains in performance to be had from write-buffering from those that can be ascribed to HTM policy. We believe that this information would facilitate sound design decisions when incorporating HTMs into parallel architectures. Anurag Negi, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Per Stenström |
ICPP | 3 |
| 2011 | ZEBRA: a data-centric, hybrid-policy hardware transactional memory designabstractHardware Transactional Memory (HTM) systems, in prior research, have either fixed policies of conflict resolution and data versioning for the entire system or allowed a degree of flexibility at the level of transactions. Unfortunately, this results in susceptibility to pathologies, lower average performance over diverse workload characteristics or high design complexity. In this work we explore a new dimension along which flexibility in policy can be introduced. Recognizing the fact that contention is more a property of data rather than that of an atomic code block, we develop an HTM system that allows selection of versioning and conflict resolution policies at the granularity of cache lines. We discover that this neat match in granularity with that of the cache coherence protocol results in a design that is very simple and yet able to track closely or exceed the performance of the best performing policy for a given workload. It also brings together the benefits of parallel commits (inherent in traditional eager HTMs) and good optimistic concurrency without deadlock avoidance mechanisms (inherent in lazy HTMs), with little increase in complexity. J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström |
ICS | 3 |
| 2011 | GLocks: Efficient Support for Highly-Contended Locks in Many-Core CMPsabstractSynchronization is of paramount importance to exploit thread-level parallelism on many-core CMPs. In these architectures, synchronization mechanisms usually rely on shared variables to coordinate multithreaded access to shared data structures thus avoiding data dependency conflicts. Lock synchronization is known to be a key limitation to performance and scalability. On the one hand, lock acquisition through busy waiting on shared variables generates additional coherence activity which interferes with applications. On the other hand, lock contention causes serialization which results in performance degradation. This paper proposes and evaluates \textit{GLocks}, a hardware-supported implementation for highly-contended locks in the context of many-core CMPs. \textit{GLocks} use a token-based message-passing protocol over a dedicated network built on state-of-the-art technology. This approach skips the memory hierarchy to provide a non-intrusive, extremely efficient and fair lock implementation with negligible impact on energy consumption or die area. A comprehensive comparison against the most efficient shared-memory-based lock implementation for a set of micro benchmarks and real applications quantifies the goodness of \textit{GLocks}. Performance results show an average reduction of 42% and 14% in execution time, an average reduction of 76% and 23% in network traffic, and also an average reduction of 78% and 28% in energy-delay$^2$ product (ED$^2$P) metric for the full CMP for the micro benchmarks and the real applications, respectively. In light of our performance results, we can conclude that \textit{GLocks} satisfy our initial working hypothesis. \textit{GLocks} minimize cache-coherence network traffic due to lock synchronization which translates into reduced power consumption and execution time. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
IPDPS | 3 |
| 2010 | EMC2: Extending Magny-Cours coherence for large-scale serversabstractThe demand of larger and more powerful high-performance shared-memory servers is growing over the last few years. To meet this need, AMD has recently launched the twelve-core Magny-Cours processors. They include a directory cache (Probe Filter) that increases the scalability of the coherence protocol applied by Opterons, based on coherent Hyper Transport interconnect (cHT). cHT limits up to 8 the number of nodes that can be addressed. Recent High Node Count HT specification overcomes this limitation. However, the 3-bit pointer used by the Probe Filter prevents Magny-Cours-based servers from being built beyond 8 nodes. In this paper, we propose and develop an external logic to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the Probe Filter. Evaluation results for up to a 32-node system show how the performance offered by our solution scales with the increment in the number of nodes, enhancing the Probe Filter effectiveness by filtering additional messages. Particularly, we reduce runtime by 47% in a 32-die system respect to the 8-die Magny-Cours system. Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato |
HiPC | 5 |
| 2010 | A G-Line-Based Network for Fast and Efficient Barrier Synchronization in Many-Core CMPsabstractBarrier synchronization in shared memory parallel machines has been widely implemented through busy-waiting on shared variables. However, typical implementations of barrier synchronization tend to produce hot-spots in terms of memory and network contention, thus creating performance bottlenecks that become markedly more pronounced as the number of cores or processors increases. To overcome such limitations, we present a novel hardware-based barrier mechanism in the context of many-core CMPs. Our proposal is based on global interconnection lines (G-lines) and the S-CSMA technique, which have been recently used to enhance a flow control mechanism (EVC) in the context of networks-on-chip. Based on this technology, we have designed a simple and scalable G-line-based network that operates independently of the main data network, and that is aimed at carrying out barrier synchronizations efficiently. In the ideal case, our design takes only 4 cycles to perform a barrier synchronization once all cores or threads have arrived at the barrier. As a proof of concept, we examine the benefits of our proposal by comparing it with one of the best software approaches (a binary combining-tree barrier). To do so, we run several kernels and scientific applications on top of the Sim-PowerCMP performance simulator that models a 32-core CMP with a 2D-mesh network configuration. Our proposal entails average reductions in terms of execution time of 68% and 21% for kernels and scientific applications, respectively. Additionally, network traffic is also lowered by 74% and 18%, respectively. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
ICPP | 3 |
| 2010 | Energy-Efficient Hardware Prefetching for CMPs Using Heterogeneous InterconnectsabstractIn the last years high performance processor designs have evolved toward Chip-Multiprocessor (CMP) architectures that implement multiple processing cores on a single die. As the number of cores inside a CMP increases, the on-chip interconnection network will have significant impact on both overall performance and power consumption as previous studies have shown. On the other hand, CMP designs are likely to be equipped with latency hiding techniques like hardware prefetching in order to reduce the negative impact on performance that, otherwise, high cache miss rates would lead to. Unfortunately, the extra number of network messages that prefetching entails can drastically increase the amount of power consumed in the interconnect. In this work, we show how to reduce the impact of prefetching techniques in terms of power (and energy) consumption in the context of tiled CMPs. Our proposal is based on the fact that the wires used in the on-chip interconnection network can be designed with varying latency, bandwidth and power characteristics. By using a heterogeneous interconnect, where low-power wires are used for dealing with prefetched lines, significant energy savings can be obtained. Detailed simulations of a 16-core CMP show that our proposal obtains improvements of up to 30% in the power consumed by the interconnect (15-23% on average) with almost negligible cost in terms of execution time (average degradation of 2%). Antonio Flores, Juan L. Aragón, Manuel E. Acacio |
PDP | 3 |
| 2010 | Characterizing Energy Consumption in Hardware Transactional Memory SystemsabstractTransactional Memory is currently being advocated as a promising alternative to lock-based synchronization because it simplifies multithreaded programming. In this way, future many-core CMP architectures may need to provide hardware support for transactional memory. On the other hand, power dissipation constitutes a first class consideration in multicore processor design. In this work, we characterize the performance and energy consumption of two well-known Hardware Transactional Memory systems that employ opposite policies for data versioning and conflict management. More specifically, we compare the Log TM-SE Eager-Eager system and a version of the Scalable TCC Lazy-Lazy system that enables parallel commits. To the best of our knowledge, this is the first characterization in terms of energy consumption of hardware transactional memory systems. To do that, we extended the GEMS simulator to estimate the energy consumed in the on-chip caches according to CACTI, and used the interconnection network energy model given by Orion 2. Results show that the energy consumption of the Eager-Eager system is 60% higher on average than in the Lazy-Lazy case, whereas performance differences between the two systems are 42% on average. Finally, we found that although on average Lazy-Lazy beats Eager-Eager there are considerable deviations in performance depending on the particular characteristics of each application. Epifanio Gaona-Ramírez, J. Rubén Titos Gil, Juan Fernández Peinador, Manuel E. Acacio |
SBAC-PAD | 4 |
| 2010 | Exploiting address compression and heterogeneous interconnects for efficient message management in tiled CMPs
Antonio Flores, Manuel E. Acacio, Juan L. Aragón |
J. Syst. Archit. | 2 |
| 2010 | A scalable organization for distributed directories
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
J. Syst. Archit. | 2 |
| 2010 | Heterogeneous Interconnects for Energy-Efficient Message Management in CMPsabstractContinuous improvements in integration scale have made major microprocessor vendors to move to designs that integrate several processing cores on the same chip. Chip multiprocessors (CMPs) constitute a good alternative to traditional monolithic designs for several reasons, among others, better levels of performance, scalability, and performance/energy ratio. On the other hand, higher clock frequencies and the increasing transistor density have revealed power dissipation and temperature as critical design issues in current and future architectures. Previous studies have shown that the interconnection network of a Chip Multiprocessor (CMP) has significant impact on both overall performance and energy consumption. Moreover, wires used in such interconnect can be designed with varying latency, bandwidth, and power characteristics. In this work, we show how messages can be efficiently managed, from the point of view of both performance and energy, in tiled CMPs using a heterogeneous interconnect. Our proposal consists of two approaches. The first is Reply Partitioning, a technique that splits replies with data into a short Partial Reply message that carries a subblock of the cache line that includes the word requested by the processor plus an Ordinary Reply with the full cache line. This technique allows all messages used to ensure coherence between the L1 caches of a CMP to be classified into two groups: critical and short, and noncritical and long. The second approach is the use of a heterogeneous interconnection network composed of low-latency wires for critical messages and low-energy wires for noncritical ones. Detailed simulations of 8 and 16-core CMPs show that our proposal obtains average savings of 7 percent in execution time and 70 percent in the Energy-Delay squared Product (ED2P) metric of the interconnect over previous works (from 24 to 30 percent average ED2P improvement for the full CMP). Additionally, the sensitivity analysis shows that although the execution time is minimized for subblocks of 16 bytes, the best choice from the point of view of the ED2P metric is the 4-byte subblock configuration with an additional improvement of 2 percent over the 16-byte one for the ED2P metric of the full CMP. Antonio Flores, Juan L. Aragón, Manuel E. Acacio |
IEEE Trans. Computers | 3 |
| 2010 | Characterizing the basic synchronization and communication operations in Dual Cell-based Blades through CellStats
José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2010 | Dealing with Transient Faults in the Interconnection Network of CMPs at the Cache Coherence LevelabstractThe importance of transient faults is predicted to grow due to current technology trends of increased scale of integration. One of the components that will be significantly affected by transient faults is the interconnection network of chip multiprocessors (CMPs). To deal efficiently with these faults and differently from other authors, we propose to use fault-tolerant cache coherence protocols that ensure the correct execution of programs when not all messages are correctly delivered. We describe the extensions made to a directory-based cache coherence protocol to provide fault tolerance and provide a modified set of token counting rules which are useful to design fault-tolerant token-based cache coherence protocols. We compare the directory-based fault-tolerant protocol with a token-based fault-tolerant one. We also show how to adjust the fault tolerance parameters to achieve the desired level of fault tolerance and measure the overhead achieved to be able to support very high fault rates. Simulation results using a set of scientific, multimedia, and commercial applications show that the fault tolerance measures have virtually no impact on execution time with respect to a non-fault-tolerant protocol. Additionally, our protocols can support very high rates of transient faults at the cost of slightly increased network traffic. Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | A Direct Coherence Protocol for Many-Core Chip MultiprocessorsabstractFuture many-core CMP designs that will integrate tens of processor cores on-chip will be constrained by area and power. Area constraints make impractical the use of a bus or a crossbar as the on-chip interconnection network, and tiled CMPs organized around a direct interconnection network will probably be the architecture of choice. Power constraints make impractical to rely on broadcasts (as, for example, Token-CMP does) or any other brute-force method for keeping cache coherence, and directory-based cache coherence protocols are currently being employed. Unfortunately, directory protocols introduce indirection to access directory information, which negatively impacts performance. In this work, we present DiCo-CMP, a novel cache coherence protocol especially suited to future many-core tiled CMP architectures. In DiCo-CMP, the task of storing up-to-date sharing information and ensuring ordered accesses for every memory block is assigned to the cache that must provide the block on a miss. Therefore, DiCo-CMP reduces the miss latency compared to a directory protocol by sending requests directly to the cache that provides the block in a cache miss. These latency reductions result in improvements in execution time of up to 6 percent, on average, over a directory protocol. In comparison with Token-CMP, our protocol only sends one request message for each cache miss, as such is able to reduce network traffic by 43 percent. Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Dealing with Traffic-Area Trade-Off in Direct Coherence Protocols for Many-Core CMPs
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
APPT | 2 |
| 2009 | Fast and Efficient Synchronization and Communication Collective Primitives for Dual Cell-Based Blades
Epifanio Gaona-Ramírez, Juan Fernández Peinador, Manuel E. Acacio |
Euro-Par | 3 |
| 2009 | Distance-aware round-robin mapping for large NUCA cachesabstractIn many-core architectures, memory blocks are commonly assigned to the banks of a NUCA cache by following a physical mapping. This mapping assigns blocks to cache banks in a round-robin fashion, thus neglecting the distance between the cores that most frequently access every block and the corresponding NUCA bank for the block. This issue impacts both cache access latency and the amount of on-chip network traffic generated. On the other hand, first-touch mapping policies, which take into account distance, can lead to an unbalanced utilization of cache banks, and consequently, to an increased number of expensive off-chip accesses. In this work, we propose the distance-aware round-robin mapping policy, an OS-managed policy which addresses the trade-off between cache access latency and number of off-chip accesses. Our policy tries to map the pages accessed by a core to its closest (local) bank, like in a first-touch policy. However, our policy also introduces an upper bound on the deviation of the distribution of memory pages among cache banks, which lessens the number of off-chip accesses. This tradeoff is addressed without requiring any extra hardware structure. We also show that the private cache indexing commonly used in many-core architectures is not the most appropriate for OS-managed distance-aware mapping policies, and propose to employ different bits for such indexing. Using GEMS simulator we show that our proposal obtains average improvements of 11% for parallel applications and 14% for multi-programmed workloads in terms of execution time, and significant reductions in network traffic, over a traditional physical mapping. Moreover, when compared to a first-touch mapping policy, our proposal improves average execution time by 5% for parallel applications and 6% for multi-programmed workloads, slightly increasing on-chip network traffic. Alberto Ros 0001, Marcelo Cintra, Manuel E. Acacio, José M. García 0001 |
HiPC | 3 |
| 2009 | Speculation-based conflict resolution in hardware transactional memoryabstractConflict management is a key design dimension of hardware transactional memory (HTM) systems, and the implementation of efficient mechanisms for detection and resolution becomes critical when conflicts are not a rare event. Current designs address this problem from two opposite perspectives, namely, lazy and eager schemes. While the former approach is based on an purely optimistic view that is not well-suited when conflicts become frequent, the latter results too pessimistic because resolves conflicts too conservatively, often limiting concurrency unnecessarily. In this paper, we present a hybrid, pseudo-optimistic scheme of conflict resolution for HTM systems that recaptures the concept of speculation to allow transactions to continue their execution past conflicting accesses. Simulation results show that our proposal is capable of combining the advantages of both classical approaches. For the STAMP transactional benchmarks, our hybrid scheme outperforms both eager and lazy systems with average reductions in execution time of 8 and 17%, respectively, and it decreases network traffic by another 17% compared to the eager policy. J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001 |
IPDPS | 2 |
| 2009 | A Parallel Implementation of the 2D Wavelet Transform Using CUDAabstractThere is a multicore platform that is currently concentrating an enormous attention due to its tremendous potential in terms of sustained performance: the NVIDIA Tesla boards. These cards intended for general-purpose computing on graphic processing units (GPGPUs) are used as data-parallel computing devices. They are based on the Computed Unified Device Architecture (CUDA) which is common to the latest NVIDIA GPUs. The bottom line is a multicore platform which provides an enormous potential performance benefit driven by a non-traditional programming model. In this paper we try to provide some insight into the peculiarities of CUDA in order to target scientific computing by means of a specific example. In particular, we show that the parallelization of the two-dimensional fast wavelet transform for the NVIDIA Tesla C870 achieves a speedup of 20.8 for an image size of 8192x8192, when compared with the fastest host-only version implementation using OpenMP and including the data transfers between main memory and device memory. Joaquín Franco, Gregorio Bernabé, Juan Fernández Peinador, Manuel E. Acacio |
PDP | 4 |
| 2008 | A fault-tolerant directory-based cache coherence protocol for CMP architecturesabstractCurrent technology trends of increased scale of integration are pushing CMOS technology into the deep-submicron domain, enabling the creation of chips with a significantly greater number of transistors but also more prone to transient failures. Hence, computer architects will have to consider reliability as a prime concern for future chip-multiprocessor designs (CMPs). Since the interconnection network of future CMPs will use a significant portion of the chip real state, it will be especially affected by transient failures. We propose to deal with this kind of failures at the level of the cache coherence protocol instead of ensuring the reliability of the network itself. Particularly, we have extended a directory-based cache coherence protocol to ensure correct program semantics even in presence of transient failures in the interconnection network. Additionally, we show that our proposal has virtually no impact on execution time with respect to a non fault-tolerant protocol, and just entails modest hardware and network traffic overhead. Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato |
DSN | 3 |
| 2008 | Fault-Tolerant Cache Coherence Protocols for CMPs: Evaluation and Trade-Offs
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato |
HiPC | 3 |
| 2008 | Directory-Based Conflict Detection in Hardware Transactional Memory
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001 |
HiPC | 2 |
| 2008 | Address Compression and Heterogeneous Interconnects for Energy-Efficient High-Performance in Tiled CMPsabstractPrevious studies have shown that the interconnection network of a Chip-Multiprocessor (CMP) has significant impact on both overall performance and energy consumption. Moreover, wires used in such interconnect can be designed with varying latency, bandwidth and power characteristics. In this work, we present a proposal for performance- and energy-efficient message management in tiled CMPs that combines both address compression with a heterogeneous interconnect. Our proposal consists of applying an address compression scheme that dynamically compresses the addresses within coherence messages allowing for a significant area slack. The arising area can be exploited for wire latency improvement by using a heterogeneous interconnection network comprised of a small set of very-low-latency wires for critical short-messages in addition to baseline wires. Detailed simulations of a 16-core CMP show that our proposal obtains average improvements of 10% in execution time and 38% in the Energy-Delay^2 Product of the interconnect. Antonio Flores, Manuel E. Acacio, Juan L. Aragón |
ICPP | 2 |
| 2008 | DiCo-CMP: Efficient cache coherency in tiled CMP architecturesabstractFuture CMP designs that will integrate tens of processor cores on-chip will be constrained by area and power. Area constraints make impractical the use of a bus or a crossbar as the on-chip interconnection network, and tiled CMPs organized around a direct interconnection network will probably be the architecture of choice. Power constraints make impractical to rely on broadcasts (as Token-CMP does) or any other brute-force method for keeping cache coherence, and directory-based cache coherence protocols are currently being employed. Unfortunately, directory protocols introduce indirection to access directory information, which negatively impacts performance. In this work, we present DiCo-CMP, a novel cache coherence protocol especially suited to future tiled CMP architectures. In DiCo- CMP the role of storing up-to-date sharing information and ensuring totally ordered accesses for every memory block is assigned to the cache that must provide the block on a miss. Therefore, DiCo-CMP reduces the miss latency compared to a directory protocol by sending coherence messages directly from the requesting caches to those that must observe them (as it would be done in brute-force protocols), and reduces the network traffic compared to Token-CMP (and consequently, power consumption in the interconnection network) by sending just one request message for each miss. Using an extended version of GEMS simulator we show that DiCo-CMP achieves improvements in execution time of up to 8% on average over a directory protocol, and reductions in terms of network traffic of up to 42% on average compared to Token-CMP. Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
IPDPS | 2 |
| 2008 | CellStats: A Tool to Evaluate the Basic Synchronization and Communication Operations of the Cell BEabstractThe Cell Broadband Engine (Cell BE) is a recent heterogeneous chip-multiprocessor (CMP) architecture jointly developed by IBM, Sony and Toshiba to offer very high performance, especially on game and multimedia applications. The significant number of processor cores that it contains (nine in its first generation), along with their heterogeneity (they are of two different types) and the variety of synchronization and communication primitives offered to programmers, make the task of developing efficient applications for the Cell BE very challenging. In this work, we present CellStats, a tool aimed at characterizing the performance of the main synchronization and communication primitives provided by the Cell BE under varying workloads. In particular, the current implementation of CellStats allows to evaluate the DMA transfer mechanism, the read-modify-write atomic operations, the mailboxes, the signals and the time taken by thread creation. As an example of application of CellStats, we present a characterization of the Cell BE incorporated into the PlayStation 3. From this characterization, we extract some recommendations that can help programmers to identify the most appropriate primitive under different assumptions. José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio |
PDP | 3 |
| 2008 | Characterization of Conflicts in Log-Based Transactional Memory (LogTM)abstractThe difficulty of multithreaded programming remains a major obstacle for programmers to fully exploit multicore chips. Transactional memory has been proposed as an abstraction capable of ameliorating the challenges of traditional lock-based parallel programming. Hardware transactional memory (HTM) systems implement the necessary mechanisms to provide transactional semantics efficiently. In order to keep hardware simple, current HTM designs apply fixed policies that aim at optimizing the most expected application behaviour, and many of these proposals explicitly assume that commits will be clearly more frequent than aborts in future transactional workloads. This paper shows that some applications developed under the TM programming model are by nature prone to experience many conflicts. As a result, aborted transactions can get to be common and may seriously hurt performance. Our characterization, performed with truly transactional benchmarks on the LogTM system, shows that certain programs composed by large transactions suffer indeed very high abort rates. Thus, if TM is to unburden developers from the programmability-performance trade-off, HTM systems must obtain good performance levels in the presence of frequent aborts, requiring more flexible policies of data versioning as well as more sophisticated recovery schemes. J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001 |
PDP | 2 |
| 2008 | Two proposals for the inclusion of directory information in the last-level private caches of glueless shared-memory multiprocessors
Alberto Ros 0001, Ricardo Fernández-Pascual, Manuel E. Acacio, José M. García 0001 |
J. Parallel Distributed Comput. | 3 |
| 2008 | An energy consumption characterization of on-chip interconnection networks for tiled CMP architectures
Antonio Flores, Juan L. Aragón, Manuel E. Acacio |
J. Supercomput. | 3 |
| 2008 | Extending the TokenCMP Cache Coherence Protocol for Low Overhead Fault Tolerance in CMP ArchitecturesabstractIt is widely accepted that transient failures will appear more frequently in chips designed in the near future due to several factors such as the increased integration scale. On the other hand, Chip-multiprocessors (CMP) that integrate several processor cores in a single chip are nowadays the best alternative to more efficient use of the increasing number of transistors that can be placed in a single die. Hence, it is necessary to design new techniques to deal with these faults to be able to build sufficiently reliable Chip Multiprocessors (CMPs). In this work, we present a coherence protocol aimed at dealing with transient failures that affect the interconnection network of a CMP, thus assuming that the network is no longer reliable. In particular, our proposal extends a token-based cache coherence protocol so that no data can be lost and no deadlock can occur due to any dropped message. Using GEMS full system simulator, we compare our proposal against TokenCMP. We show that in absence of failures our proposal does not introduce overhead in terms of increased execution time over TokenCMP. Additionally, our protocol can tolerate message loss rates much higher than those likely to be found in the real world without increasing execution time more than 15%. Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | Efficient Message Management in Tiled CMP Architectures Using a Heterogeneous Interconnection Network
Antonio Flores, Juan L. Aragón, Manuel E. Acacio |
HiPC | 3 |
| 2007 | Direct Coherence: Bringing Together Performance and Scalability in Shared-Memory Multiprocessors
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
HiPC | 2 |
| 2007 | A Low Overhead Fault Tolerant Coherence Protocol for CMP ArchitecturesabstractIt is widely accepted that transient failures will appear more frequently in chips designed in the near future due to several factors such as the increased integration scale. On the other hand, chip-multiprocessors (CMP) that integrate several processor cores in a single chip are nowadays the best alternative to more efficient use of the increasing number of transistors that can be placed in a single die. Hence, it is necessary to design new techniques to deal with these faults to be able to build sufficiently reliable chip multiprocessors (CMPs). In this work, we present a coherence protocol aimed at dealing with transient failures that affect the interconnection network of a CMP, thus assuming that the network is no longer reliable. In particular, our proposal extends a token-based cache coherence protocol so that no data can be lost and no deadlock can occur due to any dropped message. Using GEMS full system simulator, we compare our proposal against a similar protocol without fault tolerance (TOKENCMP). We show that in absence of failures our proposal does not introduce overhead in terms of increased execution time over TOKENCMP. Additionally, our protocol can tolerate message loss rates much higher than those likely to be found in the real world without increasing execution time more than 15% Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato |
HPCA | 3 |
| 2007 | An efficient implementation of a 3D wavelet transform based encoder on hyper-threading technology
Gregorio Bernabé, Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José González 0002 |
Parallel Comput. | 4 |
| 2005 | A Novel Lightweight Directory Architecture for Scalable Shared-Memory Multiprocessors
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001 |
Euro-Par | 2 |
| 2005 | Memory Subsystem Characterization in a 16-Core Snoop-Based Chip-Multiprocessor Architecture
Francisco J. Villa, Manuel E. Acacio, José M. García 0001 |
HPCC | 2 |
| 2005 | Evaluating IA-32 web servers through simics: a practical experience
Francisco J. Villa, Manuel E. Acacio, José M. García 0001 |
J. Syst. Archit. | 2 |
| 2005 | A Two-Level Directory Architecture for Highly Scalable cc-NUMA MultiprocessorsabstractOne important issue the designer of a scalable shared-memory multiprocessor must deal with is the amount of extra memory required to store the directory information. It is desirable that the directory memory overhead be kept as low as possible, and that it scales very slowly with the size of the machine. Unfortunately, current directory architectures provide scalability at the expense of performance. This work presents a scalable directory architecture that significantly reduces the size of the directory for large-scale configurations of a multiprocessor without degrading performance. First, we propose multilayer clustering as an effective approach to reduce the width of directory entries. Based on this concept, we derive three new compressed sharing codes, some of them with a space complexity of O(log/sub 2/(log/sub 2/(N))) for an N-node system. Then, we present a novel two-level directory architecture to eliminate the penalty caused by compressed directories in general. The proposed organization consists of a small full-map first-level directory (which provides precise information for the most recently referenced lines) and a compressed second-level directory (which provides in-excess information for all the lines). The proposals are evaluated based on extensive execution-driven simulations (using RSIM) of a 64-node cc-NUMA multiprocessor. Results demonstrate that a system with a two-level directory architecture achieves the same performance as a multiprocessor with a big and nonscalable full-map directory, with a very significant reduction of the memory overhead. Manuel E. Acacio, José González 0002, José M. García 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | An Architecture for High-Performance Scalable Shared-Memory Multiprocessors Exploiting On-Chip IntegrationabstractRecent technology improvements allow multiprocessor designers to put some key components inside the processor chip, such as the memory controller, the coherence hardware, and the network interface/router. In this paper, we exploit such integration scale, presenting a novel node architecture aimed at reducing the long L2 miss latencies and the memory overhead of using directories that characterize cc-NUMA machines and limit their scalability. Our proposal replaces the traditional directory with a novel three-level directory architecture, as well as it adds a small shared data cache to each of the nodes of a multiprocessor system. Due to their small size, the first-level directory and the shared data cache are integrated into the processor chip in every node, which enhances performance by saving accesses to the slower main memory. Scalability is guaranteed by having the second and third-level directories out of the processor chip and using compressed data structures. A taxonomy of the L2 misses, according to the actions performed by the directory to satisfy them, is also presented. Using execution-driven simulations, we show that significant latency reductions can be obtained by using the proposed node architecture, which translates into reductions of more than 30 percent in several cases in the application execution time. Manuel E. Acacio, José González 0002, José M. García 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | Owner prediction for accelerating cache-to-cache transfer misses in a cc-NUMA architectureabstractCache misses for which data must be obtained from a remote cache (cache-to-cache transfer misses) account for an important fraction of the total miss rate. Unfortunately, cc-NUMA designs put the access to the directory information into the critical path of 3-hop misses, which significantly penalizes them compared to SMP designs. This work studies the use of owner prediction as a means of providing cc-NUMA multiprocessors with a more efficient support for cache-to-cache transfer misses. Our proposal comprises an effective prediction scheme as well as a coherence protocol designed to support the use of prediction. Results indicate that owner prediction can significantly reduce the latency of cache-to-cache transfer misses, which translates into speed-ups on application performance up to 12%. In order to also accelerate most of those 3-hop misses that are either not predicted or mispredicted, the inclusion of a small and fast directory cache in every node is evaluated, leading to improvements up to 16% on the final performance. Manuel E. Acacio, José González 0002, José M. García 0001, José Duato |
SC | 1 |
| 2002 | MPI-Delphi: an MPI implementation for visual programming environments and heterogeneous computing
Manuel E. Acacio, Óscar Cánovas Reverte, José M. García 0001, Pedro E. López-de-Teruel |
Future Gener. Comput. Syst. | 1 |
| 2001 | A New Scalable Directory Architecture for Large-Scale MultiprocessorsabstractThe memory overhead introduced by directories constitutes a major hurdle in the scalability of cc-NUMA architectures, which makes the shared-memory paradigm unfeasible for very large-scale systems. This work is focused on improving the scalability of shared-memory multiprocessors by significantly reducing the size of the directory. We propose multilayer clustering as an effective approach to reduce the directory-entry width. Detailed evaluation for 64 processors shows that using this approach we can drastically reduce the memory overhead, while suffering a performance degradation we similar to previous compressed schemes (such as Coarse Vector). In addition, a novel two-level directory architecture is proposed in order to eliminate the penalty caused by these compressed directories. This organization consists of a small Full-Map first-level directory (which provides precise information for the most recently referenced lines) and a compressed second-level directory (which provides in-excess information). Results show that a system with this directory architecture can achieve the same performance as a multiprocessor with a big and non-scalable Full-Map directory with a very significant reduction of the memory overhead. Manuel E. Acacio, José González 0002, José M. García 0001, José Duato |
HPCA | 1 |