EDBT 2026 Demo / reviewers in the wild / expert
Ram Rangan
dblp:67/394
· DBLP profile ↗
18ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0003-4191-4151ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
GPUs and heterogeneous computing · 44% Processor architecture and microarchitecture · 26% Memory systems · 19% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Computer graphics and multimedia
2 papers |
Rendering · 100% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › data locality
cache locality |
0.6 | 1 | 2022 | Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022 |
GPUs and heterogeneous computing
control flow divergence |
0.6 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
GPUs and heterogeneous computing
GPU microarchitecture |
0.6 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
Processor architecture and microarchitecture
latency hiding |
0.6 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
Processor architecture and microarchitecture › pipelining
pipeline stall |
0.6 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
Memory systems › cache
texture cache |
0.6 | 1 | 2022 | Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022 |
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling |
0.6 | 1 | 2022 | Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.6 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.4 | 1 | 2020 | Zeroploit: Exploiting Zero Valued Operands in Interactive Gaming Applications · ACM Trans. Archit. Code Optim. 2020 |
Rendering
ray tracing |
0.2 | 1 | 2022 | GPU Subwarp Interleaving · HPCA 2022 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 3 | 2008 | Support for High-Frequency Streaming in CMPs · MICRO 2006 Performance scalability of decoupled software pipelining · ACM Trans. Archit. Code Optim. 2008 Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005 |
Processor architecture and microarchitecture
instruction set architecture |
0.1 | 1 | 2009 | Lightweight predication support for out of order processors · HPCA 2009 |
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor |
0.1 | 1 | 2009 | Lightweight predication support for out of order processors · HPCA 2009 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.1 | 1 | 2009 | Lightweight predication support for out of order processors · HPCA 2009 |
Parallel and multicore computing › parallel computing › parallel communication
inter-thread communication |
0.1 | 1 | 2006 | Support for High-Frequency Streaming in CMPs · MICRO 2006 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.1 | 1 | 2005 | Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005 |
Hardware reliability and fault tolerance › soft errors
architectural vulnerability factor |
0.1 | 1 | 2005 | Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005 |
Parallel and multicore computing › parallelization strategies
automatic thread extraction |
0.1 | 1 | 2005 | Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005 |
Electronic design automation › hardware verification and test
fault detection |
0.1 | 1 | 2005 | Software-controlled fault tolerance · ACM Trans. Archit. Code Optim. 2005 |
Hardware reliability and fault tolerance › reliability analysis
lifetime analysis |
0.1 | 1 | 2005 | Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2005 | Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005 |
Hardware reliability and fault tolerance › soft errors
soft error mitigation |
0.1 | 1 | 2005 | Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005 |
Parallel and multicore computing
thread-level parallelism |
0.1 | 1 | 2005 | Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005 |
Hardware reliability and fault tolerance › error detection
transient fault detection |
0.1 | 1 | 2005 | Design and Evaluation of Hybrid Fault-Detection Systems · ISCA 2005 |
Systems and software security
information flow control |
0.0 | 1 | 2004 | RIFLE: An Architectural Framework for User-Centric Information-Flow Security · MICRO 2004 |
Processor architecture and microarchitecture
hardware-assisted security |
0.0 | 1 | 2004 | RIFLE: An Architectural Framework for User-Centric Information-Flow Security · MICRO 2004 |
Compilers and program optimization
compilation target |
0.0 | 1 | 2009 | Lightweight predication support for out of order processors · HPCA 2009 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.0 | 1 | 2008 | Performance scalability of decoupled software pipelining · ACM Trans. Archit. Code Optim. 2008 |
Memory systems › memory architecture
memory system support |
0.0 | 1 | 2006 | Support for High-Frequency Streaming in CMPs · MICRO 2006 |
Electronic design automation › circuit simulation
reliability simulation |
0.0 | 1 | 2005 | Design and Evaluation of Hybrid Fault-Detection Systems · ISCA 2005 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.3profile-guided code optimization · 1.3fast-slow versioning · 1.3backward/forward slice specialization · 1.3microbenchmarking · 1.1software techniques for cache locality · 0.6predicated mutually exclusive groups · 0.2hammock predication · 0.2performance scalability analysis · 0.1streaming-aware memory enhancement · 0.1compiler implementation · 0.1binary translation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | GPU Subwarp InterleavingabstractRaytracing applications have naturally high thread divergence, low warp occupancy and are limited by memory latency. In this paper, we present an architectural enhancement called Subwarp Interleaving that exploits thread divergence to hide pipeline stalls in divergent sections of low warp occupancy workloads. Subwarp Interleaving allows for fine-grained interleaved execution of diverged paths within a warp with the goal of increasing hardware utilization and reducing warp latency. However, notwithstanding the promise shown by early microbenchmark studies and an average performance upside of 6.3% (up to 20%) on a simulator across a suite of raytradng application traces, the Subwarp Interleaving design feature has shortcomings that preclude its near-term implementation. This paper introduces the reader to the challenges of raytradng and discusses a novel micro-architectural approach that, on paper, addresses many of the challenges. A thorough analysis of the idea on a production simulator reveals that the high-level motivating statistics are optimistic, and second-order effects, along with other architectural sharp edges, limit the idea’s potential. We identify Subwarp Interleaving’s primary limiters for an NVIDIA Tbring-like architecture, and we outline the conditions under which the approach could be more effective. Sana Damani, Mark Stephenson, Ram Rangan, Daniel R. Johnson, Rishkul Kulkami, Stephen W. Keckler |
HPCA | 3 |
| 2022 | Locality-Aware CTA Scheduling for Gaming ApplicationsabstractThe compute work rasterizer or the GigaThread Engine of a modern NVIDIA GPU focuses on maximizing compute work occupancy across all streaming multiprocessors in a GPU while retaining design simplicity. In this article, we identify the operational aspects of the GigaThread Engine that help it meet those goals but also lead to less-than-ideal cache locality for texture accesses in 2D compute shaders, which are an important optimization target for gaming applications. We develop three software techniques, namely LargeCTAs , Swizzle , and Agents , to show that it is possible to effectively exploit the texture data working set overlap intrinsic to 2D compute shaders. We evaluate these techniques on gaming applications across two generations of NVIDIA GPUs, RTX 2080 and RTX 3080, and find that they are effective on both GPUs. We find that the bandwidth savings from all our software techniques on RTX 2080 is much higher than the bandwidth savings on baseline execution from inter-generational cache capacity increase going from RTX 2080 to RTX 3080. Our best-performing technique, Agents , records up to a 4.7% average full-frame speedup by reducing bandwidth demand of targeted shaders at the L1-L2 and L2-DRAM interfaces by 23% and 32%, respectively, on the latest generation RTX 3080. These results acutely highlight the sensitivity of cache locality to compute work rasterization order and the importance of locality-aware cooperative thread array scheduling for gaming applications. Aditya Ukarande, Suryakant Patidar, Ram Rangan |
ACM Trans. Archit. Code Optim. | 3 |
| 2021 | PGZ: automatic zero-value code specializationabstractIn prior work we proposed Zeroploit, a transform that duplicates code, specializes one path assuming certain key program operands, called versioning variables, are zero, and leaves the other path unspecialized. Dynamically, depending on the versioning variable’s value, either the specialized fast path or the default slow path will execute. We evaluated Zeroploit with hand-optimized codes in that work. Mark Stephenson, Ram Rangan |
CC | 2 |
| 2021 | Cooperative Profile Guided OptimizationsabstractAbstract Existing feedback‐driven optimization frameworks are not suitable for video games, which tend to push the limits of performance of gaming platforms and have real‐time constraints that preclude all but the simplest execution profiling. While Profile Guided Optimization (PGO) is a well‐established optimization approach, existing PGO techniques are ill‐suited for games for a number of reasons, particularly because heavyweight profiling makes interactive applications unresponsive. Adaptive optimization frameworks continually collect metrics that guide code specialization optimizations during program execution but have similarly high overheads. We emulate a system, which we callCooperative PGO, in which the gaming platform collectspiecemealprofiles by sampling in both time and space during actual gameplay across many users; stitches the piecemeal profiles together statistically; and creates policies to guide future gameplay. We introduce a three‐level hierarchical profiler that is well‐suited to graphics APIs, that commonly operates with no overhead and occasionally introduces an average overhead of less than 0.5% during periods of active profiling. This paper examines the practicality ofCooperative PGOusing three PGOs as case studies. A PGO that exploits likely zeros is particularly effective, achieving an average speedup of 5%, with a maximum speedup of 15%, over a highly‐tuned baseline. Mark Stephenson, Ram Rangan, Stephen W. Keckler |
Comput. Graph. Forum | 2 |
| 2020 | Zeroploit: Exploiting Zero Valued Operands in Interactive Gaming ApplicationsabstractIn this article, we first characterize register operand value locality in shader programs of modern gaming applications and observe that there is a high likelihood of one of the register operands of several multiply, logical-and, and similar operations being zero, dynamically. We provide intuition, examples, and a quantitative characterization for how zeros originate dynamically in these programs. Next, we show that this dynamic behavior can be gainfully exploited with a profile-guided code optimization called Zeroploit that transforms targeted code regions into a zero-(value-)specialized fast path and a default slow path. The fast path benefits from zero-specialization in two ways, namely: (a) the backward slice of the other operand of a given multiply or logical-and can be skipped dynamically, provided the only use of that other operand is in the given instruction, and (b) the forward slice of instructions originating at the given instruction can be zero-specialized, potentially triggering further backward slice specializations from operations of that forward slice as well. Such specialization helps the fast path avoid redundant dynamic computations as well as memory fetches, while the fast-slow versioning transform helps preserve functional correctness. With an offline value profiler and manually optimized shader programs, we demonstrate that Zeroploit is able to achieve an average speedup of 35.8% for targeted shader programs, amounting to an average frame-rate speedup of 2.8% across a collection of modern gaming applications on an NVIDIA® GeForce RTX™ 2080 GPU. Ram Rangan, Mark Stephenson, Aditya Ukarande, Shyam Murthy, Virat Agarwal, Marc Blackstein |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Mesoscale performance simulation of multicore processor systems
Peter Altevogt, Tibor Kiss, Michael Kistler, Ram Rangan |
Softw. Syst. Model. | 4 |
| 2010 | Statistically regulating program behavior via mainstream computingabstractWe introduce mainstream computing, a collaborative system that dynamically checks a program--via runtime assertion checks--to ensure that it is running according to expectation. Rather than enforcing strict, statically-defined assertions, our system allows users to run with a set of assertions that are statistically guaranteed to fail at a rate bounded by a user-defined probability, pfail. For example, a user can request a set of assertions that will fail at most 0.5% of the times the application is invoked. Users who believe their usage of an application is mainstream can use relatively large settings for pfail. Higher values of pfail provide stricter regulation of the application which likely enhances security, but will also inhibit some legitimate program behaviors; in contrast, program behavior is unregulated when pfail = 0, leaving the user vulnerable to attack. We show that our prototype is able to detect denial of service attacks, integer overflows, frees of uninitialized memory, boundary violations, and an injection attack. In addition we perform experiments with a mainstream computing system designed to protect against soft errors. Mark Stephenson, Ram Rangan, Emmanuel Yashchin, Eric Van Hensbergen |
CGO | 2 |
| 2009 | Lightweight predication support for out of order processorsabstractThe benefits of Out of Order (OOO) processing are well known, as is the effectiveness of predicated execution for unpredictable control flow. However, as previous research has demonstrated, these techniques are at odds with one another. One common approach to reconciling their differences is to simplify the form of predication supported by the architecture. For instance, the only form of predication supported by modern OOO processors is a simple conditional move. We argue that it is the simplicity of conditional move that has allowed its widespread adoption, but we also show that this simplicity compromises its effectiveness as a compilation target. In this paper, we introduce a generalized form of hammock predication - called predicated mutually exclusive groups - that requires few modifications to an existing processor pipeline, yet presents the compiler with abundant predication opportunities. In comparison to non-predicated code running on an aggressively clocked baseline system, our technique achieves an 8% speedup averaged across three important benchmark suites. Mark Stephenson, Lixin Zhang 0002, Ram Rangan |
HPCA | 3 |
| 2008 | Spice: speculative parallel iteration chunk executionabstractThe recent trend in the processor industry of packing multiple processor cores in a chip has increased the importance of automatic techniques for extracting thread level parallelism. A promising approach for extracting thread level parallelism in general purpose applications is to apply memory alias or value speculation to break dependences amongst threads and executes them concurrently. Easwaran Raman, Neil Vachharajani, Ram Rangan, David I. August |
CGO | 3 |
| 2008 | Performance scalability of decoupled software pipeliningabstractAny successful solution to using multicore processors to scale general-purpose program performance will have to contend with rising intercore communication costs while exposing coarse-grained parallelism. Recently proposed pipelined multithreading (PMT) techniques have been demonstrated to have general-purpose applicability and are also able to effectively tolerate inter-core latencies through pipelined interthread communication. These desirable properties make PMT techniques strong candidates for program parallelization on current and future multicore processors and understanding their performance characteristics is critical to their deployment. To that end, this paper evaluates the performance scalability of a general-purpose PMT technique called decoupled software pipelining (DSWP) and presents a thorough analysis of the communication bottlenecks that must be overcome for optimal DSWP scalability. Ram Rangan, Neil Vachharajani, Guilherme Ottoni, David I. August |
ACM Trans. Archit. Code Optim. | 1 |
| 2007 | Speculative Decoupled Software Pipelining
Neil Vachharajani, Ram Rangan, Easwaran Raman, Matthew J. Bridges, Guilherme Ottoni, David I. August |
PACT | 2 |
| 2006 | Support for High-Frequency Streaming in CMPsabstractAs the industry moves toward larger-scale chip multiprocessors, the need to parallelize applications grows. High inter-thread communication delays, exacerbated by over-stressed high-latency memory subsystems and ever-increasing wire delays, require parallelization techniques to create partially or fully independent threads to improve performance. Unfortunately, developers and compilers alike often fail to find sufficient independent work of this kind. Recently proposed pipelined streaming techniques have shown significant promise for both manual and automatic parallelization. These techniques have wide-scale applicability because they embrace inter-thread dependences (albeit acyclic dependences) and tolerate long-latency communication of these dependences. This paper addresses the lack of architectural support for this type of concurrency, which has blocked its adoption and hindered related language and compiler research. We observe that both manual and automatic techniques create high-frequency streaming threads, with communication occurring every 5 to 20 instructions. Even while easily tolerating inter-thread transit delays, high-frequency communication makes thread performance very sensitive to intrathread delays from the repeated execution of the communication operations. Using this observation, we define the design-space and evaluate several mechanisms to find a better trade-off between performance and operating system, hardware, and design costs. From this, we find a light-weight streaming-aware enhancement to conventional memory subsystems that doubles the speed of these codes and is within 2% of the best-performing, but heavy-weight, hardware solution. Ram Rangan, Neil Vachharajani, Adam Stoler, Guilherme Ottoni, David I. August, George Z. N. Cai |
MICRO | 1 |
| 2005 | SWIFT: Software Implemented Fault ToleranceabstractTo improve performance and reduce power, processor designers employ advances that shrink feature sizes, lower voltage levels, reduce noise margins, and increase clock rates. However, these advances make processors more susceptible to transient faults that can affect correctness. While reliable systems typically employ hardware techniques to address soft-errors, software techniques can provide a lower-cost and more flexible alternative. This paper presents a novel, software-only, transient-fault-detection technique, called SWIFT. SWIFT efficiently manages redundancy by reclaiming unused instruction-level resources present during the execution of most programs. SWIFT also provides a high level of protection and performance with an enhanced control-flow checking mechanism. We evaluate an implementation of SWIFT on an Itanium 2 which demonstrates exceptional fault coverage with a reasonable performance cost. Compared to the best known single-threaded approach utilizing an ECC memory system, SWIFT demonstrates a 51% average speedup. George A. Reis, Neil Vachharajani, Ram Rangan, David I. August |
CGO | 4 |
| 2005 | Computing Architectural Vulnerability Factors for Address-Based StructuresabstractProcessor designers require estimates of the architectural vulnerability factor (AVF) of on-chip structures to make accurate soft error rate estimates. AVF is the fraction of faults from alpha particle and neutron strikes that result in user-visible errors. This paper shows how to use a performance model to calculate the AVF of address-based structures, using a data cache, a data translation buffer, and a store buffer as examples. We describe how to perform a detailed breakdown of lifetime components (e.g., fill-to-read, read-to-evict) of bits in these structures into ACE (required for architecturally correct execution), un-ACE (unnecessary for ACE), and unknown components. This lifetime analysis produces best estimate AVFs for these three structures' data arrays of 6%, 36%, and 4%, respectively. We then present a new technique, hamming-distance-one analysis, and show that it predicts surprisingly low best estimate AVFs of 0.41%, 3%, and 7.7% for the structures' tag arrays. Finally, using our lifetime analysis framework, we show how two AVF reduction techniques - periodic flushing and incremental scrubbing - can reduce the AVF by converting ACE lifetime components into un-ACE without affecting performance significantly. Arijit Biswas, Paul Racunas, Razvan Cheveresan, Joel S. Emer, Shubhendu S. Mukherjee, Ram Rangan |
ISCA | 6 |
| 2005 | Design and Evaluation of Hybrid Fault-Detection SystemsabstractAs chip densities and clock rates increase, processors are becoming more susceptible to transient faults that can affect program correctness. Up to now, system designers have primarily considered hardware-only and software-only fault-detection mechanisms to identify and mitigate the deleterious effects of transient faults. These two fault-detection systems, however, are extremes in the design space, representing sharp trade-offs between hardware cost, reliability, and performance. In this paper, we identify hybrid hardware/software fault-detection mechanisms as promising alternatives to hardware-only and software-only systems. These hybrid systems offer designers more options to fit their reliability needs within their hardware and performance budgets. We propose and evaluate CRAFT, a suite of three such hybrid techniques, to illustrate the potential of the hybrid approach. For fair, quantitative comparisons among hardware, software, and hybrid systems, we introduce a new metric, mean work to failure, which is able to compare systems for which machine instructions do not represent a constant unit of work. Additionally, we present a new simulation framework which rapidly assesses reliability and does not depend on manual identification of failure modes. Our evaluation illustrates that CRAFT, and hybrid techniques in general, offer attractive options in the fault-detection design space. George A. Reis, Neil Vachharajani, Ram Rangan, David I. August, Shubhendu S. Mukherjee |
ISCA | 4 |
| 2005 | Automatic Thread Extraction with Decoupled Software PipeliningabstractUntil recently, a steadily rising clock rate and other uniprocessor micro architectural improvements could be relied upon to consistently deliver increasing performance for a wide range of applications. Current difficulties in maintaining this trend have lead microprocessor manufacturers to add value by incorporating multiple processors on a chip. Unfortunately, since decades of compiler research have not succeeded in delivering automatic threading for prevalent code properties, this approach demonstrates no improvement for a large class of existing codes. To find useful work for chip multiprocessors, we propose an automatic approach to thread extraction, called decoupled software pipelining (DSWP). DSWP exploits the finegrained pipeline parallelism lurking in most applications to extract long-running, concurrently executing threads. Use of the nonspeculative and truly decoupled threads produced by DSWP can increase execution efficiency and provide significant latency tolerance, mitigating design complexity by reducing intercore communication and per-core resource requirements. Using our initial fully automatic compiler implementation and a validated processor model, we prove the concept by demonstrating significant gains for dual-core chip multiprocessor models running a variety of codes. We then explore simple opportunities missed by our initial compiler implementation which suggest a promising future for this approach. Guilherme Ottoni, Ram Rangan, Adam Stoler, David I. August |
MICRO | 2 |
| 2005 | Software-controlled fault toleranceabstractTraditional fault-tolerance techniques typically utilize resources ineffectively because they cannot adapt to the changing reliability and performance demands of a system. This paper proposes software-controlled fault tolerance, a concept allowing designers and users to tailor their performance and reliability for each situation. Several software-controllable fault-detection techniques are then presented: SWIFT, a software-only technique, and CRAFT, a suite of hybrid hardware/software techniques. Finally, the paper introduces PROFiT, a technique which adjusts the level of protection and performance at fine granularities through software control. When coupled with software-controllable techniques like SWIFT and CRAFT, PROFiT offers attractive and novel reliability options. George A. Reis, Neil Vachharajani, Ram Rangan, David I. August, Shubhendu S. Mukherjee |
ACM Trans. Archit. Code Optim. | 4 |
| 2004 | RIFLE: An Architectural Framework for User-Centric Information-Flow SecurityabstractEven as modern computing systems allow the manipulation and distribution of massive amounts of information, users of these systems are unable to manage the confidentiality of their data in a practical fashion. Conventional access control security mechanisms cannot prevent the illegitimate use of privileged data once access is granted. For example, information provided by a user during an online purchase may be covertly delivered to malicious third parties by an untrustworthy web browser. Existing information-flow security mechanisms do provide this assurance, but only for programmer-specified policies enforced during program development as a static analysis on special-purpose type-safe languages. Not only are these techniques not applicable to many commonly used programs, but they leave the user with no defense against malicious programmers or altered binaries. In this paper, we propose RIFLE, a runtime information-flow security system designed from the user's perspective. By addressing information-flow security using architectural support, RIFLE gives users a practical way to enforce their own information-flow security policy on all programs. We prove that, contrary to statements in the literature, run-time systems like RIFLE are no less secure than existing language-based techniques. Using a model of the architectural framework and a binary translator, we demonstrate RIFLE's correctness and illustrate that the performance cost is reasonable. Neil Vachharajani, Matthew J. Bridges, Ram Rangan, Guilherme Ottoni, Jason A. Blome, George A. Reis, Manish Vachharajani, David I. August |
MICRO | 4 |