Ravi Rajwar

dblp:98/6810 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Memory systems · 34% Processor architecture and microarchitecture · 26% Cloud and datacenter computing · 22%
Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 46% Transaction processing and concurrency control · 46% Database system architecture and tuning · 7%
Software engineering, system software, and programming languages
5 papers
Concurrent programming · 58% Compilers and program optimization · 24% Operating systems · 18%

Topics — the 30 heaviest of 45, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
DRAM
0.712023
Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023
Memory systems
tiered memory
0.712023
Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023
Cloud and datacenter computing
warehouse-scale computing
0.712023
Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023
Parallel and multicore computing
transactional memory
0.332013
Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013
Virtualizing Transactional Memory · ISCA 2005
Transactional lock-free execution of lock-based programs · ASPLOS 2002
Cloud and datacenter computing › resource management
datacenter memory management
0.212023
Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale · ASPLOS (3) 2023
Parallel and multicore computing
synchronization
0.222013
Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013
Improving the Throughput of Synchronization by Insertion of Delays · HPCA 2000
Indexing and storage engines
b+-tree
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Transaction processing and concurrency control › transactional memory
hardware transactional memory
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Indexing and storage engines
in-memory index
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Transaction processing and concurrency control
synchronization
0.212014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Parallel and multicore computing › transactional memory
hardware transactional memory
0.212013
Performance evaluation of Intel® transactional synchronization extensions for high-performance computing · SC 2013
Processor architecture and microarchitecture › out-of-order execution
instruction window
0.132004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
Continual flow pipelines · ASPLOS 2004
Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window Processors · MICRO 2003
Processor architecture and microarchitecture
speculative execution
0.122007
Hardware atomicity for reliable software speculation · ISCA 2007
Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001
Concurrent programming
synchronization
0.122005
Virtualizing Transactional Memory · ISCA 2005
Transactional lock-free execution of lock-based programs · ASPLOS 2002
Processor architecture and microarchitecture › checkpoint-based microarchitecture
checkpoint processing and recovery
0.122004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window Processors · MICRO 2003
Database system architecture and tuning
main-memory database
0.112014
Improving in-memory database index performance with Intel® Transactional Synchronization Extensions · HPCA 2014
Processor architecture and microarchitecture
branch prediction
0.122004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
An Architectural Evaluation of Java TPC-W · HPCA 2001
Concurrent programming
transactional memory
0.112005
Virtualizing Transactional Memory · ISCA 2005
Processor architecture and microarchitecture
chip multiprocessor
0.112005
The Impact of Performance Asymmetry in Emerging Multicore Architectures · ISCA 2005
Processor architecture and microarchitecture › load/store queue
load-store queue management
0.112005
Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005
Processor architecture and microarchitecture
memory latency tolerance
0.112005
Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering
0.112005
Scalable Load and Store Processing in Latency Tolerant Processors · ISCA 2005
Processor architecture and microarchitecture › branch prediction
branch misprediction recovery
0.012004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
Processor architecture and microarchitecture › pipelining
continual flow pipeline
0.012004
Continual flow pipelines · ASPLOS 2004
Processor architecture and microarchitecture › pipelining
pipeline design
0.012004
Continual flow pipelines · ASPLOS 2004
Processor architecture and microarchitecture
register file
0.012004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
Processor architecture and microarchitecture › memory system microarchitecture
store-load forwarding
0.012004
An analysis of a resource efficient checkpoint architecture · ACM Trans. Archit. Code Optim. 2004
Concurrent programming › synchronization
locking
0.012001
Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001
Parallel and multicore computing › transactional memory
lock elision
0.012001
Speculative lock elision: enabling highly concurrent multithreaded execution · MICRO 2001
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.012001
An Architectural Evaluation of Java TPC-W · HPCA 2001

Methods — techniques the papers use, named apart from their topics

application-transparent tiering · 0.7Intel TSX · 0.4hardware transactional memory · 0.3hardware prototype evaluation · 0.1simulation · 0.1store redo log · 0.1secondary load buffer · 0.1cache-based forwarding · 0.1selective checkpointing · 0.0hierarchical store queue · 0.0timestamps · 0.0rollback recovery · 0.0cache-based conflict detection · 0.0
YearPublicationVenuePosition
2023 Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-Scale
abstract
Fast DRAM increasingly dominates infrastructure spend in large scale computing environments and this trend will likely worsen without an architectural shift. The cost of deployed memory can be reduced by replacing part of the conventional DRAM with lower cost albeit slower memory media, thus creating a tiered memory system where both tiers are directly addressable and cached. But, this poses numerous challenges in a highly multi-tenant warehouse-scale computing setting. The diversity and scale of its applications motivates an application-transparent solution in the general case, adaptable to specific workload demands.
Padmapriya Duraisamy, Scott Hare, Ravi Rajwar, David E. Culler, Zhiyi Xu, Jianing Fan, Chris Kennelly, Bill McCloskey, Danijela Mijailovic, Brian Morris, Chiranjit Mukherjee, Jingliang Ren, Greg Thelen, Carlos Villavieja, Parthasarathy Ranganathan, Amin Vahdat
ASPLOS (3)4
2015 Specialized Evolution of the General Purpose CPU
Ravi Rajwar, Martin Dixon, Ronak Singhal
CIDR1
2014 Improving in-memory database index performance with Intel® Transactional Synchronization Extensions
abstract
The increasing number of cores every generation poses challenges for high-performance in-memory database systems. While these systems use sophisticated high-level algorithms to partition a query or run multiple queries in parallel, they also utilize low-level synchronization mechanisms to synchronize access to internal database data structures. Developers often spend significant development and verification effort to improve concurrency in the presence of such synchronization. The Intel®Transactional Synchronization Extensions (Intel®TSX) in the 4th Generation Core™ Processors enable hardware to dynamically determine whether threads actually need to synchronize even in the presence of conservatively used synchronization. This paper evaluates the effectiveness of such hardware support in a commercial database. We focus on two index implementations: a B+Tree Index and the Delta Storage Index used in the SAP HANA®database system. We demonstrate that such support can improve performance of database data structures such as index trees and presents a compelling opportunity for the development of simpler, scalable, and easy-to-verify algorithms.
Tomas Karnagel, Roman Dementiev, Ravi Rajwar, Konrad Lai, Thomas Legler, Benjamin Schlegel, Wolfgang Lehner
HPCA3
2013 Performance evaluation of Intel® transactional synchronization extensions for high-performance computing
abstract
Intel has recently introduced Intel® Transactional Synchronization Extensions (Intel® TSX) in the Intel 4th Generation Core™ Processors. With Intel TSX, a processor can dynamically determine whether threads need to serialize through lock-protected critical sections. In this paper, we evaluate the first hardware implementation of Intel TSX using a set of high-performance computing (HPC) workloads, and demonstrate that applying Intel TSX to these workloads can provide significant performance improvements. On a set of real-world HPC workloads, applying Intel TSX provides an average speedup of 1.41x. When applied to a parallel user-level TCP/IP stack, Intel TSX provides 1.31x average bandwidth improvement on network intensive applications. We also demonstrate the ease with which we were able to apply Intel TSX to the various workloads.
Richard M. Yoo, Christopher J. Hughes, Konrad Lai, Ravi Rajwar
SC4
2012 In search of parallel dimensions
abstract
Performance matters. But how we improve it is changing. Historically, transparent hardware improvements would mean software just ran faster. That may not necessarily be true in the future. To continue the pace of innovation, the future will need to be increasingly parallel--nvolving parallelism across data, threads, cores, and nodes. This talk will explore some of the dimensions of parallelism, and the opportunities and challenges they pose.
Ravi Rajwar
SPAA1
2007 Hardware atomicity for reliable software speculation
abstract
Speculative compiler optimizations are effective in improving both single-thread performance and reducing power consumption, but their implementation introduces significant complexity, which can limit their adoption, limit their optimization scope, and negatively impact the reliability of the compilers that implement them. To eliminate much of this complexity, as well as increase the effectiveness of these optimizations, we propose that microprocessors provide architecturally-visible hardware primitives for atomic execution. These primitives provide to the compiler the ability to optimize the program's hot path in isolation, allowing the use of non-speculative formulations of optimization passes to perform speculative optimizations. Atomic execution guarantees that if a speculation invariant does not hold, the speculative updates are discarded, the register state is restored, and control is transferred to a non-speculative version of the code, thereby relieving the compiler from the responsibility of generating compensation code.
Naveen Neelakantam, Ravi Rajwar, Suresh Srinivas, Uma Srinivasan 0003, Craig B. Zilles
ISCA2
2007 Transactional memory and the birthday paradox
abstract
Many word-based Software TransactionalMemory systems (STMs) have been proposed using tagless ownership tables, where read and write permissions are granted at the granularity of all addresses that map to a given ownership table entry. This optimization to reduce overhead potentially results in false conflicts. Using address traces from a multithreaded program, we demonstrate that the frequency of these false conflicts grows superlinearly with both the TM data footprint and concurrency and that increasing the size of the ownership table results in only a sub-linear reduction in conflict rate. These somewhat surprising relationships have a theoretical foundation that is also responsible for the (naively) unintuitive statistical result generally referred to as the "Birthday Paradox." We present an analytical model based on random population of an ownership table by concurrently executing transactions that correctly predicts the trends in measured data. These results call into question the viability of such an optimization that can undermine the scalability and concurrency claims of software transactional memory.
Craig B. Zilles, Ravi Rajwar
SPAA2
2005 The Impact of Performance Asymmetry in Emerging Multicore Architectures
abstract
Performance asymmetry in multicore architectures arises when individual cores have different performance. Building such multicore processors is desirable because many simple cores together provide high parallel performance while a few complex cores ensure high serial performance. However, application developers typically assume computational cores provide equal performance, and performance asymmetry breaks this assumption. This paper is concerned with the behavior of commercial applications running on performance asymmetric systems. We present the first study investigating the impact of performance asymmetry on a wide range of commercial applications using a hardware prototype. We quantify the impact of asymmetry on an application's performance variance when run multiple times, and the impact on the application's scalability. Performance asymmetry adversely affects behavior of many workloads. We study ways to eliminate these effects. In addition to asymmetry-aware operating system kernels, the application often itself needs to be aware of performance asymmetry for stable and scalable performance.
Saisanthosh Balakrishnan, Ravi Rajwar, Michael Upton, Konrad Lai
ISCA2
2005 Scalable Load and Store Processing in Latency Tolerant Processors
abstract
Memory latency tolerant architectures support thousands of in-flight instructions without scaling cycle-critical processor resources, and thousands of useful instructions can complete in parallel with a miss to memory. These architectures however require large queues to track all loads and stores executed while a miss is pending. Hierarchical designs alleviate cycle time impact of these structures but the CAM and search functions required to enforce memory ordering and provide data forwarding place high demand on area and power. We present new load-store processing algorithms for latency tolerant architectures. We augment primary load and store queues with secondary buffers. The secondary load buffer is a set associative structure, similar to a cache. The secondary store buffer, the Store Redo Log, is a first-in first-out structure recording the program order of all stores completed in parallel with a miss, and has no CAM and search functions. Instead of the secondary store queue, a cache provides temporary forwarding. The SRL enforces memory ordering by ensuring memory updates occur in program order once the miss returns. The new algorithms eliminate the CAM and search functions in the secondary load and store buffers, and remove fundamental sources of complexity, power, and area inefficiency in load/store processing. The new organization, while being area and power efficient, is competitive in performance compared to hierarchical designs.
Amit Gandhi, Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan, Konrad Lai
ISCA3
2005 Virtualizing Transactional Memory
abstract
Writing concurrent programs is difficult because of the complexity of ensuring proper synchronization. Conventional lock-based synchronization suffers from well-known limitations, so researchers have considered nonblocking transactions as an alternative. Recent hardware proposals have demonstrated how transactions can achieve high performance while not suffering limitations of lock-based mechanisms. However, current hardware proposals require programmers to be aware of platform-specific resource limitations such as buffer sizes, scheduling quanta, as well as events such as page faults, and process migrations. If the transactional model is to gain wide acceptance, hardware support for transactions must be virtualized to hide these limitations in much the same way that virtual memory shields the programmer from platform-specific limitations of physical memory. This paper proposes virtual transactional memory (VTM), a user-transparent system that shields the programmer from various platform-specific resource limitations. VTM maintains the performance advantage of hardware transactions, incurs low overhead in time, and has modest costs in hardware support. While many system-level challenges remain, VTM takes a step toward making transactional models more widely acceptable.
Ravi Rajwar, Maurice Herlihy, Konrad Lai
ISCA1
2004 Continual flow pipelines
abstract
Increased integration in the form of multiple processor cores on a single die, relatively constant die sizes, shrinking power envelopes, and emerging applications create a new challenge for processor architects. How to build a processor that provides high single-thread performance and enables multiple of these to be placed on the same die for high throughput while dynamically adapting for future applications? Conventional approaches for high single-thread performance rely on large and complex cores to sustain a large instruction window for memory tolerance, making them unsuitable for multi-core chips. We present Continual Flow Pipelines (CFP) as a new non-blocking processor pipeline architecture that achieves the performance of a large instruction window without requiring cycle-critical structures such as the scheduler and register file to be large. We show that to achieve benefits of a large instruction window, inefficiencies in management of both the scheduler and register file must be addressed, and we propose a unified solution. The non-blocking property of CFP keeps key processor structures affecting cycle time and power (scheduler, register file), and die size (second level cache) small. The memory latency-tolerant CFP core allows multiple cores on a single die while outperforming current processor cores for single-thread applications.
Srikanth T. Srinivasan, Ravi Rajwar, Haitham Akkary, Amit Gandhi, Michael Upton
ASPLOS2
2004 An analysis of a resource efficient checkpoint architecture
abstract
Large instruction window processors achieve high performance by exposing large amounts of instruction level parallelism. However, accessing large hardware structures typically required to buffer and process such instruction window sizes significantly degrade the cycle time. This paper proposes a novel checkpoint processing and recovery (CPR) microarchitecture, and shows how to implement a large instruction window processor without requiring large structures thus permitting a high clock frequency.We focus on four critical aspects of a microarchitecture: (1) scheduling instructions, (2) recovering from branch mispredicts, (3) buffering a large number of stores and forwarding data from stores to any dependent load, and (4) reclaiming physical registers. While scheduling window size is important, we show the performance of large instruction windows to be more sensitive to the other three design issues. Our CPR proposal incorporates novel microarchitectural schemes for addressing these design issues---a selective checkpoint mechanism for recovering from mispredicts, a hierarchical store queue organization for fast store-load forwarding, and an effective algorithm for aggressive physical register reclamation. Our proposals allow a processor to realize performance gains due to instruction windows of thousands of instructions without requiring large cycle-critical hardware structures.
Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan
ACM Trans. Archit. Code Optim.2
2003 Inferential queueing and speculative push for reducing critical communication latencies
abstract
Communication latencies within critical sections constitute a major bottleneck in some classes of emerging parallel workloads. In this paper, we argue for the use of Inferentially Queued Locks (IQLs) [31], not just for efficient synchronization but also for reducing communication latencies, and we propose a novel mechanism, Speculative Push (SP), aimed at reducing these communication latencies. With IQLs, the processor infers the existence, and limits, of a critical section from the use of synchronization instructions and joins a queue of lock requestors. The SP mechanism extracts information about program structure by observing IQLs. SP allows the cache controller, responding to a request for a cache line that likely includes a lock variable, to predict the data sets the requestor will modify within the associated critical section. The controller then pushes these lines from its own cache to the target cache, as well as writing them to memory. Overlapping the protected data transfer with that of the lock can substantially reduce the communication latencies within critical sections. By pushing data in exclusive state, the mechanism can collapse a read-modify-write sequences within a critical section into a single local cache access. The write-back to memory allows the receiving cache to ignore the push. Neither mechanism requires any programmer or compiler support nor any instruction set changes. Our experiments demonstrate that IQLs and SP can improve performance of applications employing frequent synchronization.
Ravi Rajwar, Alain Kägi, James R. Goodman
ICS1
2003 Checkpoint Processing and Recovery: Towards Scalable Large Instruction Window Processors
abstract
Large instruction window processors achieve high performance by exposing large amounts of instruction level parallelism. However, accessing large hardware structures typically required to buffer and process such instruction window sizes significantly degrade the cycle time. This paper proposes a checkpoint processing and recovery (CPR) microarchitecture, and shows how to implement a large instruction window processor without requiring large structures thus permitting a high clock frequency. We focus of four critical aspects of a microarchitecture: 1) scheduling instructions; 2) recovering from branch mispredicts; 3) buffering a large number of stores and forwarding data from stores to any dependent load; and 4) reclaiming physical registers. While scheduling window size is important, we show the performance of large instruction windows to be more sensitive to the other three design issues. Our CPR proposal incorporates novel microarchitecture scheme for addressing these design issues-a selective checkpoint mechanism for recovering from mispredicts, a hierarchical store queue organization for fast store-load forwarding, and an effective algorithm for aggressive physical register reclamation. Our proposals allow a processor to realize performance gains due to instruction windows of thousands of instructions without requiring large cycle-critical hardware structures.
Haitham Akkary, Ravi Rajwar, Srikanth T. Srinivasan
MICRO2
2002 Transactional lock-free execution of lock-based programs
abstract
This paper is motivated by the difficulty in writing correct high-performance programs. Writing shared-memory multi-threaded programs imposes a complex trade-off between programming ease and performance, largely due to subtleties in coordinating access to shared data. To ensure correctness programmers often rely on conservative locking at the expense of performance. The resulting serialization of threads is a performance bottleneck. Locks also interact poorly with thread scheduling and faults, resulting in poor system performance.We seek to improve multithreaded programming trade-offs by providing architectural support for optimistic lock-free execution. In a lock-free execution, shared objects are never locked when accessed by various threads. We propose Transactional Lock Removal (TLR) and show how a program that uses lock-based synchronization can be executed by the hardware in a lock-free manner, even in the presence of conflicts, without programmer support or software changes. TLR uses timestamps for conflict resolution, modest hardware, and features already present in many modern computer systems.TLR's benefits include improved programmability, stability, and performance. Programmers can obtain benefits of lock-free data structures, such as non-blocking behavior and wait-freedom, while using lock-protected critical sections for writing programs.
Ravi Rajwar, James R. Goodman
ASPLOS1
2001 An Architectural Evaluation of Java TPC-W
abstract
The use of the Java programming language for implementing server-side application logic is increasingly in popularity yet there is very little known about the architectural requirements of this emerging commercial workload. We present a detailed characterization of the Transaction Processing Council's TPC-W web benchmark, implemented in Java. The TPC-W benchmark is designed to exercise the web server and transaction processing system of a typical e-commerce web site. We have implemented TPC-W as a collection of Java servlets, and present an architectural study detailing the memory system and branch predictor behavior of the workload. We also evaluate the effectiveness of a coarse-grained multithreaded processor at increasing system throughput using TPC-W and other commercial workloads. We measure system throughput improvements from 8% to 41% for a two context processor, and 12% to 60% for a four context uniprocessor over a single-threaded uniprocessor despite decreased branch prediction accuracy and cache hit rates.
Harold W. Cain, Ravi Rajwar, Morris Marden, Mikko H. Lipasti
HPCA2
2001 Speculative lock elision: enabling highly concurrent multithreaded execution
abstract
Serialization of threads due to critical sections is a fundamental bottleneck to achieving high performance in multithreaded programs. Dynamically, such serialization may be unnecessary because these critical sections could have safely executed concurrently without locks. Current processors cannot fully exploit such parallelism because they do not have mechanisms to dynamically detect such false inter-thread dependences. We propose Speculative Lock Elision (SLE), a novel micro-architectural technique to remove dynamically unnecessary lock-induced serialization and enable highly concurrent multithreaded execution. The key insight is that locks do not always have to be acquired for a correct execution. Synchronization instructions are predicted as being unnecessary and elided. This allows multiple threads to concurrently execute critical sections protected by the same lock. Misspeculation due to inter-thread data conflicts is detected using existing cache mechanisms and rollback is used for recovery. Successful speculative elision is validated and committed without acquiring the lock. SLE can be implemented entirely in microarchitecture without instruction set support and without system-level modifications, is transparent to programmers, and requires only trivial additional hardware support. SLE can provide programmers a fast path to writing correct high-performance multithreaded programs.
Ravi Rajwar, James R. Goodman
MICRO1
2000 Improving the Throughput of Synchronization by Insertion of Delays
abstract
Efficiency of synchronization mechanisms can limit the parallel performance of many shared-memory applications. In addition, the ever increasing performance gap between processor and interprocessor communication may further compromise the scalability of these primitives. Ideally, synchronization primitives should provide high performance under both high and low contention without requiring substantial programmer effort and software support. QOLR has been shown to offer substantial speedups and to outperform other synchronization primitives consistently, but at the cost of software support and protocol complexity. This paper proposes the use of speculation and delays to implement a purely hardware-based queueing mechanism called Implicit QOLB. Making use of the pervasiveness of the Load-Linked/Store-Conditional primitives, we present a series of hardware mechanisms to optimize performance for sharing patterns exhibited by locks and associated data. The mechanisms do not require any change to existing software or instruction sets. IQOLB sits alongside the cache-coherence protocol and guides the decisions the protocol makes with respect to lock (and associated data) transfers. Preliminary evaluations indicate that IQOLB may perform as well as, if not better than, QOLB without the additional software and protocol complexity.
Ravi Rajwar, Alain Kägi, James R. Goodman
HPCA1