EDBT 2026 Demo / reviewers in the wild / expert
Ravi Rajwar
dblp:98/6810
· DBLP profile ↗
18ranked-venue papers
7as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Memory systems · 34% Processor architecture and microarchitecture · 26% Cloud and datacenter computing · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Indexing and storage engines · 46% Transaction processing and concurrency control · 46% Database system architecture and tuning · 7% | |
| Software engineering, system software, and programming languages
5 papers |
Concurrent programming · 58% Compilers and program optimization · 24% Operating systems · 18% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
DRAM |
0.7 | 1 | 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023 |
Memory systems
tiered memory |
0.7 | 1 | 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023 |
Cloud and datacenter computing
warehouse-scale computing |
0.7 | 1 | 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023 |
Parallel and multicore computing
transactional memory |
0.3 | 3 | 2013 | Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013 Virtualizing Transactional Memory · ISCA 2005 Transactional lock-free execution of lock-based programs · ASPLOS 2002 |
Cloud and datacenter computing › resource management
datacenter memory management |
0.2 | 1 | 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023 |
Parallel and multicore computing
synchronization |
0.2 | 2 | 2013 | Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013 Improving the Throughput of Synchronization by Insertion of Delays · HPCA 2000 |
Indexing and storage engines
b+-tree |
0.2 | 1 | 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014 |
Transaction processing and concurrency control › transactional memory
hardware transactional memory |
0.2 | 1 | 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014 |
Indexing and storage engines
in-memory index |
0.2 | 1 | 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014 |
Transaction processing and concurrency control
synchronization |
0.2 | 1 | 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.2 | 1 | 2013 | Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.1 | 3 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 Continual flow pipelines · ASPLOS 2004 Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window Processors · MICRO 2003 |
Processor architecture and microarchitecture
speculative execution |
0.1 | 2 | 2007 | Hardware atomicity for reliable software speculation · ISCA 2007 Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001 |
Concurrent programming
synchronization |
0.1 | 2 | 2005 | Virtualizing Transactional Memory · ISCA 2005 Transactional lock-free execution of lock-based programs · ASPLOS 2002 |
Processor architecture and microarchitecture › checkpoint-based microarchitecture
checkpoint processing and recovery |
0.1 | 2 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window Processors · MICRO 2003 |
Database system architecture and tuning
main-memory database |
0.1 | 1 | 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 2 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 An Architectural Evaluation of Java TPC-W · HPCA 2001 |
Concurrent programming
transactional memory |
0.1 | 1 | 2005 | Virtualizing Transactional Memory · ISCA 2005 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2005 | The Impact of Performance Asymmetry in Emerging Multicore Architectures · ISCA 2005 |
Processor architecture and microarchitecture › load/store queue
load-store queue management |
0.1 | 1 | 2005 | Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005 |
Processor architecture and microarchitecture
memory latency tolerance |
0.1 | 1 | 2005 | Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005 |
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering |
0.1 | 1 | 2005 | Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005 |
Processor architecture and microarchitecture › branch prediction
branch misprediction recovery |
0.0 | 1 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 |
Processor architecture and microarchitecture › pipelining
continual flow pipeline |
0.0 | 1 | 2004 | Continual flow pipelines · ASPLOS 2004 |
Processor architecture and microarchitecture › pipelining
pipeline design |
0.0 | 1 | 2004 | Continual flow pipelines · ASPLOS 2004 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 |
Processor architecture and microarchitecture › memory system microarchitecture
store-load forwarding |
0.0 | 1 | 2004 | An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004 |
Concurrent programming › synchronization
locking |
0.0 | 1 | 2001 | Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001 |
Parallel and multicore computing › transactional memory
lock elision |
0.0 | 1 | 2001 | Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 1 | 2001 | An Architectural Evaluation of Java TPC-W · HPCA 2001 |
Methods — techniques the papers use, named apart from their topics
application-transparent tiering · 0.7Intel TSX · 0.4hardware transactional memory · 0.3hardware prototype evaluation · 0.1simulation · 0.1store redo log · 0.1secondary load buffer · 0.1cache-based forwarding · 0.1selective checkpointing · 0.0hierarchical store queue · 0.0timestamps · 0.0rollback recovery · 0.0cache-based conflict detection · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScaleabstractFast DRAM increasingly dominates infrastructure spend in large scale computing environments and this trend will likely worsen without an architectural shift. The cost of deployed memory can be reduced by replacing part of the conventional DRAM with lower cost albeit slower memory media, thus creating a tiered memory system where both tiers are directly addressable and cached. But, this poses numerous challenges in a highly multi-tenant warehouse-scale computing setting. The diversity and scale of its applications motivates an application-transparent solution in the general case, adaptable to specific workload demands. Padmapriya Duraisamy, Scott Hare, Ravi Rajwar, David E. Culler, Zhiyi Xu, Jianing Fan, Chris Kennelly, Bill McCloskey, Danijela Mijailovic, Brian Morris, Chiranjit Mukherjee, Jingliang Ren, Greg Thelen, Carlos Villavieja, Parthasarathy Ranganathan, Amin Vahdat |
ASPLOS (3) | 4 |
| 2015 | Specialized Evolution of the General Purpose CPU
Ravi Rajwar, Martin Dixon, Ronak Singhal |
CIDR | 1 |
| 2014 | Improving in-memory database index performance with Intel® Transactional Synchronization ExtensionsabstractThe increasing number of cores every generation poses challenges for high-performance in-memory database systems. While these systems use sophisticated high-level algorithms to partition a query or run multiple queries in parallel, they also utilize low-level synchronization mechanisms to synchronize access to internal database data structures. Developers often spend significant development and verification effort to improve concurrency in the presence of such synchronization. The Intel®Transactional Synchronization Extensions (Intel®TSX) in the 4th Generation Core™ Processors enable hardware to dynamically determine whether threads actually need to synchronize even in the presence of conservatively used synchronization. This paper evaluates the effectiveness of such hardware support in a commercial database. We focus on two index implementations: a B+Tree Index and the Delta Storage Index used in the SAP HANA®database system. We demonstrate that such support can improve performance of database data structures such as index trees and presents a compelling opportunity for the development of simpler, scalable, and easy-to-verify algorithms. Tomas Karnagel, Roman Dementiev, Ravi Rajwar, Konrad Lai, Thomas Legler, Benjamin Schlegel, Wolfgang Lehner |
HPCA | 3 |
| 2013 | Performance evaluation of Intel® transactional synchronization extensions for high-performance computingabstractIntel has recently introduced Intel® Transactional Synchronization Extensions (Intel® TSX) in the Intel 4th Generation Core™ Processors. With Intel TSX, a processor can dynamically determine whether threads need to serialize through lock-protected critical sections. In this paper, we evaluate the first hardware implementation of Intel TSX using a set of high-performance computing (HPC) workloads, and demonstrate that applying Intel TSX to these workloads can provide significant performance improvements. On a set of real-world HPC workloads, applying Intel TSX provides an average speedup of 1.41x. When applied to a parallel user-level TCP/IP stack, Intel TSX provides 1.31x average bandwidth improvement on network intensive applications. We also demonstrate the ease with which we were able to apply Intel TSX to the various workloads. Richard M. Yoo, Christopher J. Hughes, Konrad Lai, Ravi Rajwar |
SC | 4 |
| 2012 | In search of parallel dimensionsabstractPerformance matters. But how we improve it is changing. Historically, transparent hardware improvements would mean software just ran faster. That may not necessarily be true in the future. To continue the pace of innovation, the future will need to be increasingly parallel--nvolving parallelism across data, threads, cores, and nodes. This talk will explore some of the dimensions of parallelism, and the opportunities and challenges they pose. Ravi Rajwar |
SPAA | 1 |
| 2007 | Hardware atomicity for reliable software speculationabstractSpeculative compiler optimizations are effective in improving both single-thread performance and reducing power consumption, but their implementation introduces significant complexity, which can limit their adoption, limit their optimization scope, and negatively impact the reliability of the compilers that implement them. To eliminate much of this complexity, as well as increase the effectiveness of these optimizations, we propose that microprocessors provide architecturally-visible hardware primitives for atomic execution. These primitives provide to the compiler the ability to optimize the program's hot path in isolation, allowing the use of non-speculative formulations of optimization passes to perform speculative optimizations. Atomic execution guarantees that if a speculation invariant does not hold, the speculative updates are discarded, the register state is restored, and control is transferred to a non-speculative version of the code, thereby relieving the compiler from the responsibility of generating compensation code. Naveen Neelakantam, Ravi Rajwar, Suresh Srinivas, Uma Srinivasan 0003, Craig B. Zilles |
ISCA | 2 |
| 2007 | Transactional memory and the birthday paradoxabstractMany word-based Software TransactionalMemory systems (STMs) have been proposed using tagless ownership tables, where read and write permissions are granted at the granularity of all addresses that map to a given ownership table entry. This optimization to reduce overhead potentially results in false conflicts. Using address traces from a multithreaded program, we demonstrate that the frequency of these false conflicts grows superlinearly with both the TM data footprint and concurrency and that increasing the size of the ownership table results in only a sub-linear reduction in conflict rate. These somewhat surprising relationships have a theoretical foundation that is also responsible for the (naively) unintuitive statistical result generally referred to as the "Birthday Paradox." We present an analytical model based on random population of an ownership table by concurrently executing transactions that correctly predicts the trends in measured data. These results call into question the viability of such an optimization that can undermine the scalability and concurrency claims of software transactional memory. Craig B. Zilles, Ravi Rajwar |
SPAA | 2 |
| 2005 | The Impact of Performance Asymmetry in Emerging Multicore ArchitecturesabstractPerformance asymmetry in multicore architectures arises when individual cores have different performance. Building such multicore processors is desirable because many simple cores together provide high parallel performance while a few complex cores ensure high serial performance. However, application developers typically assume computational cores provide equal performance, and performance asymmetry breaks this assumption. This paper is concerned with the behavior of commercial applications running on performance asymmetric systems. We present the first study investigating the impact of performance asymmetry on a wide range of commercial applications using a hardware prototype. We quantify the impact of asymmetry on an application's performance variance when run multiple times, and the impact on the application's scalability. Performance asymmetry adversely affects behavior of many workloads. We study ways to eliminate these effects. In addition to asymmetry-aware operating system kernels, the application often itself needs to be aware of performance asymmetry for stable and scalable performance. Saisanthosh Balakrishnan, Ravi Rajwar, Michael Upton, Konrad Lai |
ISCA | 2 |
| 2005 | Scalable Load and Store Processing in Latency Tolerant ProcessorsabstractMemory latency tolerant architectures support thousands of in-flight instructions without scaling cycle-critical processor resources, and thousands of useful instructions can complete in parallel with a miss to memory. These architectures however require large queues to track all loads and stores executed while a miss is pending. Hierarchical designs alleviate cycle time impact of these structures but the CAM and search functions required to enforce memory ordering and provide data forwarding place high demand on area and power. We present new load-store processing algorithms for latency tolerant architectures. We augment primary load and store queues with secondary buffers. The secondary load buffer is a set associative structure, similar to a cache. The secondary store buffer, the Store Redo Log, is a first-in first-out structure recording the program order of all stores completed in parallel with a miss, and has no CAM and search functions. Instead of the secondary store queue, a cache provides temporary forwarding. The SRL enforces memory ordering by ensuring memory updates occur in program order once the miss returns. The new algorithms eliminate the CAM and search functions in the secondary load and store buffers, and remove fundamental sources of complexity, power, and area inefficiency in load/store processing. The new organization, while being area and power efficient, is competitive in performance compared to hierarchical designs. Amit Gandhi, Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan, Konrad Lai |
ISCA | 3 |
| 2005 | Virtualizing Transactional MemoryabstractWriting concurrent programs is difficult because of the complexity of ensuring proper synchronization. Conventional lock-based synchronization suffers from well-known limitations, so researchers have considered nonblocking transactions as an alternative. Recent hardware proposals have demonstrated how transactions can achieve high performance while not suffering limitations of lock-based mechanisms. However, current hardware proposals require programmers to be aware of platform-specific resource limitations such as buffer sizes, scheduling quanta, as well as events such as page faults, and process migrations. If the transactional model is to gain wide acceptance, hardware support for transactions must be virtualized to hide these limitations in much the same way that virtual memory shields the programmer from platform-specific limitations of physical memory. This paper proposes virtual transactional memory (VTM), a user-transparent system that shields the programmer from various platform-specific resource limitations. VTM maintains the performance advantage of hardware transactions, incurs low overhead in time, and has modest costs in hardware support. While many system-level challenges remain, VTM takes a step toward making transactional models more widely acceptable. Ravi Rajwar, Maurice Herlihy, Konrad Lai |
ISCA | 1 |
| 2004 | Continual flow pipelinesabstractIncreased integration in the form of multiple processor cores on a single die, relatively constant die sizes, shrinking power envelopes, and emerging applications create a new challenge for processor architects. How to build a processor that provides high single-thread performance and enables multiple of these to be placed on the same die for high throughput while dynamically adapting for future applications? Conventional approaches for high single-thread performance rely on large and complex cores to sustain a large instruction window for memory tolerance, making them unsuitable for multi-core chips. We present Continual Flow Pipelines (CFP) as a new non-blocking processor pipeline architecture that achieves the performance of a large instruction window without requiring cycle-critical structures such as the scheduler and register file to be large. We show that to achieve benefits of a large instruction window, inefficiencies in management of both the scheduler and register file must be addressed, and we propose a unified solution. The non-blocking property of CFP keeps key processor structures affecting cycle time and power (scheduler, register file), and die size (second level cache) small. The memory latency-tolerant CFP core allows multiple cores on a single die while outperforming current processor cores for single-thread applications. Srikanth T. Srinivasan, Ravi Rajwar, Haitham Akkary, Amit Gandhi, Michael Upton |
ASPLOS | 2 |
| 2004 | An analysis of a resource efficient checkpoint architectureabstractLarge instruction window processors achieve high performance by exposing large amounts of instruction level parallelism. However, accessing large hardware structures typically required to buffer and process such instruction window sizes significantly degrade the cycle time. This paper proposes a novel checkpoint processing and recovery (CPR) microarchitecture, and shows how to implement a large instruction window processor without requiring large structures thus permitting a high clock frequency.We focus on four critical aspects of a microarchitecture: (1) scheduling instructions, (2) recovering from branch mispredicts, (3) buffering a large number of stores and forwarding data from stores to any dependent load, and (4) reclaiming physical registers. While scheduling window size is important, we show the performance of large instruction windows to be more sensitive to the other three design issues. Our CPR proposal incorporates novel microarchitectural schemes for addressing these design issues---a selective checkpoint mechanism for recovering from mispredicts, a hierarchical store queue organization for fast store-load forwarding, and an effective algorithm for aggressive physical register reclamation. Our proposals allow a processor to realize performance gains due to instruction windows of thousands of instructions without requiring large cycle-critical hardware structures. Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan |
ACM Trans. Archit. Code Optim. | 2 |
| 2003 | Inferential queueing and speculative push for reducing critical communication latenciesabstractCommunication latencies within critical sections constitute a major bottleneck in some classes of emerging parallel workloads. In this paper, we argue for the use of Inferentially Queued Locks (IQLs) [31], not just for efficient synchronization but also for reducing communication latencies, and we propose a novel mechanism, Speculative Push (SP), aimed at reducing these communication latencies. With IQLs, the processor infers the existence, and limits, of a critical section from the use of synchronization instructions and joins a queue of lock requestors. The SP mechanism extracts information about program structure by observing IQLs. SP allows the cache controller, responding to a request for a cache line that likely includes a lock variable, to predict the data sets the requestor will modify within the associated critical section. The controller then pushes these lines from its own cache to the target cache, as well as writing them to memory. Overlapping the protected data transfer with that of the lock can substantially reduce the communication latencies within critical sections. By pushing data in exclusive state, the mechanism can collapse a read-modify-write sequences within a critical section into a single local cache access. The write-back to memory allows the receiving cache to ignore the push. Neither mechanism requires any programmer or compiler support nor any instruction set changes. Our experiments demonstrate that IQLs and SP can improve performance of applications employing frequent synchronization. Ravi Rajwar, Alain Kägi, James R. Goodman |
ICS | 1 |
| 2003 | Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window ProcessorsabstractLarge instruction window processors achieve high performance by exposing large amounts of instruction level parallelism. However, accessing large hardware structures typically required to buffer and process such instruction window sizes significantly degrade the cycle time. This paper proposes a checkpoint processing and recovery (CPR) microarchitecture, and shows how to implement a large instruction window processor without requiring large structures thus permitting a high clock frequency. We focus of four critical aspects of a microarchitecture: 1) scheduling instructions; 2) recovering from branch mispredicts; 3) buffering a large number of stores and forwarding data from stores to any dependent load; and 4) reclaiming physical registers. While scheduling window size is important, we show the performance of large instruction windows to be more sensitive to the other three design issues. Our CPR proposal incorporates novel microarchitecture scheme for addressing these design issues-a selective checkpoint mechanism for recovering from mispredicts, a hierarchical store queue organization for fast store-load forwarding, and an effective algorithm for aggressive physical register reclamation. Our proposals allow a processor to realize performance gains due to instruction windows of thousands of instructions without requiring large cycle-critical hardware structures. Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan |
MICRO | 2 |
| 2002 | Transactional lock-free execution of lock-based programsabstractThis paper is motivated by the difficulty in writing correct high-performance programs. Writing shared-memory multi-threaded programs imposes a complex trade-off between programming ease and performance, largely due to subtleties in coordinating access to shared data. To ensure correctness programmers often rely on conservative locking at the expense of performance. The resulting serialization of threads is a performance bottleneck. Locks also interact poorly with thread scheduling and faults, resulting in poor system performance.We seek to improve multithreaded programming trade-offs by providing architectural support for optimistic lock-free execution. In a lock-free execution, shared objects are never locked when accessed by various threads. We propose Transactional Lock Removal (TLR) and show how a program that uses lock-based synchronization can be executed by the hardware in a lock-free manner, even in the presence of conflicts, without programmer support or software changes. TLR uses timestamps for conflict resolution, modest hardware, and features already present in many modern computer systems.TLR's benefits include improved programmability, stability, and performance. Programmers can obtain benefits of lock-free data structures, such as non-blocking behavior and wait-freedom, while using lock-protected critical sections for writing programs. Ravi Rajwar, James R. Goodman |
ASPLOS | 1 |
| 2001 | An Architectural Evaluation of Java TPC-WabstractThe use of the Java programming language for implementing server-side application logic is increasingly in popularity yet there is very little known about the architectural requirements of this emerging commercial workload. We present a detailed characterization of the Transaction Processing Council's TPC-W web benchmark, implemented in Java. The TPC-W benchmark is designed to exercise the web server and transaction processing system of a typical e-commerce web site. We have implemented TPC-W as a collection of Java servlets, and present an architectural study detailing the memory system and branch predictor behavior of the workload. We also evaluate the effectiveness of a coarse-grained multithreaded processor at increasing system throughput using TPC-W and other commercial workloads. We measure system throughput improvements from 8% to 41% for a two context processor, and 12% to 60% for a four context uniprocessor over a single-threaded uniprocessor despite decreased branch prediction accuracy and cache hit rates. Harold W. Cain, Ravi Rajwar, Morris Marden, Mikko H. Lipasti |
HPCA | 2 |
| 2001 | Speculative lock elision: enabling highly concurrent multithreaded executionabstractSerialization of threads due to critical sections is a fundamental bottleneck to achieving high performance in multithreaded programs. Dynamically, such serialization may be unnecessary because these critical sections could have safely executed concurrently without locks. Current processors cannot fully exploit such parallelism because they do not have mechanisms to dynamically detect such false inter-thread dependences. We propose Speculative Lock Elision (SLE), a novel micro-architectural technique to remove dynamically unnecessary lock-induced serialization and enable highly concurrent multithreaded execution. The key insight is that locks do not always have to be acquired for a correct execution. Synchronization instructions are predicted as being unnecessary and elided. This allows multiple threads to concurrently execute critical sections protected by the same lock. Misspeculation due to inter-thread data conflicts is detected using existing cache mechanisms and rollback is used for recovery. Successful speculative elision is validated and committed without acquiring the lock. SLE can be implemented entirely in microarchitecture without instruction set support and without system-level modifications, is transparent to programmers, and requires only trivial additional hardware support. SLE can provide programmers a fast path to writing correct high-performance multithreaded programs. Ravi Rajwar, James R. Goodman |
MICRO | 1 |
| 2000 | Improving the Throughput of Synchronization by Insertion of DelaysabstractEfficiency of synchronization mechanisms can limit the parallel performance of many shared-memory applications. In addition, the ever increasing performance gap between processor and interprocessor communication may further compromise the scalability of these primitives. Ideally, synchronization primitives should provide high performance under both high and low contention without requiring substantial programmer effort and software support. QOLR has been shown to offer substantial speedups and to outperform other synchronization primitives consistently, but at the cost of software support and protocol complexity. This paper proposes the use of speculation and delays to implement a purely hardware-based queueing mechanism called Implicit QOLB. Making use of the pervasiveness of the Load-Linked/Store-Conditional primitives, we present a series of hardware mechanisms to optimize performance for sharing patterns exhibited by locks and associated data. The mechanisms do not require any change to existing software or instruction sets. IQOLB sits alongside the cache-coherence protocol and guides the decisions the protocol makes with respect to lock (and associated data) transfers. Preliminary evaluations indicate that IQOLB may perform as well as, if not better than, QOLB without the additional software and protocol complexity. Ravi Rajwar, Alain Kägi, James R. Goodman |
HPCA | 1 |