Ram Rangan

dblp:67/394 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0003-4191-4151ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
GPUs and heterogeneous computing · 44% Processor architecture and microarchitecture · 26% Memory systems · 19%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Computer graphics and multimedia
2 papers
Rendering · 100%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › data locality
cache locality
0.612022
Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022
GPUs and heterogeneous computing
control flow divergence
0.612022
GPU Subwarp Interleaving · HPCA 2022
GPUs and heterogeneous computing
GPU microarchitecture
0.612022
GPU Subwarp Interleaving · HPCA 2022
Processor architecture and microarchitecture
latency hiding
0.612022
GPU Subwarp Interleaving · HPCA 2022
Processor architecture and microarchitecture › pipelining
pipeline stall
0.612022
GPU Subwarp Interleaving · HPCA 2022
Memory systems › cache
texture cache
0.612022
Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling
0.612022
Locality-Aware CTA Scheduling for Gaming Applications · ACM Trans. Archit. Code Optim. 2022
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.612022
GPU Subwarp Interleaving · HPCA 2022
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.412020
Zeroploit: Exploiting Zero Valued Operands in Interactive Gaming Applications · ACM Trans. Archit. Code Optim. 2020
Rendering
ray tracing
0.212022
GPU Subwarp Interleaving · HPCA 2022
Processor architecture and microarchitecture
chip multiprocessor
0.132008
Support for High-Frequency Streaming in CMPs · MICRO 2006
Performance scalability of decoupled software pipelining · ACM Trans. Archit. Code Optim. 2008
Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005
Processor architecture and microarchitecture
instruction set architecture
0.112009
Lightweight predication support for out of order processors · HPCA 2009
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor
0.112009
Lightweight predication support for out of order processors · HPCA 2009
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.112009
Lightweight predication support for out of order processors · HPCA 2009
Parallel and multicore computing › parallel computing › parallel communication
inter-thread communication
0.112006
Support for High-Frequency Streaming in CMPs · MICRO 2006
Compilers and program optimization › instruction scheduling
software pipelining
0.112005
Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005
Hardware reliability and fault tolerance › soft errors
architectural vulnerability factor
0.112005
Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005
Parallel and multicore computing › parallelization strategies
automatic thread extraction
0.112005
Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005
Electronic design automation › hardware verification and test
fault detection
0.112005
Software-controlled fault tolerance · ACM Trans. Archit. Code Optim. 2005
Hardware reliability and fault tolerance › reliability analysis
lifetime analysis
0.112005
Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005
Hardware reliability and fault tolerance
soft errors
0.112005
Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005
Hardware reliability and fault tolerance › soft errors
soft error mitigation
0.112005
Computing Architectural Vulnerability Factors for Address-Based Structures · ISCA 2005
Parallel and multicore computing
thread-level parallelism
0.112005
Automatic Thread Extraction with Decoupled Software Pipelining · MICRO 2005
Hardware reliability and fault tolerance › error detection
transient fault detection
0.112005
Design and Evaluation of Hybrid Fault-Detection Systems · ISCA 2005
Systems and software security
information flow control
0.012004
RIFLE: An Architectural Framework for User-Centric Information-Flow Security · MICRO 2004
Processor architecture and microarchitecture
hardware-assisted security
0.012004
RIFLE: An Architectural Framework for User-Centric Information-Flow Security · MICRO 2004
Compilers and program optimization
compilation target
0.012009
Lightweight predication support for out of order processors · HPCA 2009
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication
0.012008
Performance scalability of decoupled software pipelining · ACM Trans. Archit. Code Optim. 2008
Memory systems › memory architecture
memory system support
0.012006
Support for High-Frequency Streaming in CMPs · MICRO 2006
Electronic design automation › circuit simulation
reliability simulation
0.012005
Design and Evaluation of Hybrid Fault-Detection Systems · ISCA 2005

Methods — techniques the papers use, named apart from their topics

simulation · 1.3profile-guided code optimization · 1.3fast-slow versioning · 1.3backward/forward slice specialization · 1.3microbenchmarking · 1.1software techniques for cache locality · 0.6predicated mutually exclusive groups · 0.2hammock predication · 0.2performance scalability analysis · 0.1streaming-aware memory enhancement · 0.1compiler implementation · 0.1binary translation · 0.0
YearPublicationVenuePosition
2022 GPU Subwarp Interleaving
abstract
Raytracing applications have naturally high thread divergence, low warp occupancy and are limited by memory latency. In this paper, we present an architectural enhancement called Subwarp Interleaving that exploits thread divergence to hide pipeline stalls in divergent sections of low warp occupancy workloads. Subwarp Interleaving allows for fine-grained interleaved execution of diverged paths within a warp with the goal of increasing hardware utilization and reducing warp latency. However, notwithstanding the promise shown by early microbenchmark studies and an average performance upside of 6.3% (up to 20%) on a simulator across a suite of raytradng application traces, the Subwarp Interleaving design feature has shortcomings that preclude its near-term implementation. This paper introduces the reader to the challenges of raytradng and discusses a novel micro-architectural approach that, on paper, addresses many of the challenges. A thorough analysis of the idea on a production simulator reveals that the high-level motivating statistics are optimistic, and second-order effects, along with other architectural sharp edges, limit the idea’s potential. We identify Subwarp Interleaving’s primary limiters for an NVIDIA Tbring-like architecture, and we outline the conditions under which the approach could be more effective.
Sana Damani, Mark Stephenson, Ram Rangan, Daniel R. Johnson, Rishkul Kulkami, Stephen W. Keckler
HPCA3
2022 Locality-Aware CTA Scheduling for Gaming Applications
abstract
The compute work rasterizer or the GigaThread Engine of a modern NVIDIA GPU focuses on maximizing compute work occupancy across all streaming multiprocessors in a GPU while retaining design simplicity. In this article, we identify the operational aspects of the GigaThread Engine that help it meet those goals but also lead to less-than-ideal cache locality for texture accesses in 2D compute shaders, which are an important optimization target for gaming applications. We develop three software techniques, namely LargeCTAs , Swizzle , and Agents , to show that it is possible to effectively exploit the texture data working set overlap intrinsic to 2D compute shaders. We evaluate these techniques on gaming applications across two generations of NVIDIA GPUs, RTX 2080 and RTX 3080, and find that they are effective on both GPUs. We find that the bandwidth savings from all our software techniques on RTX 2080 is much higher than the bandwidth savings on baseline execution from inter-generational cache capacity increase going from RTX 2080 to RTX 3080. Our best-performing technique, Agents , records up to a 4.7% average full-frame speedup by reducing bandwidth demand of targeted shaders at the L1-L2 and L2-DRAM interfaces by 23% and 32%, respectively, on the latest generation RTX 3080. These results acutely highlight the sensitivity of cache locality to compute work rasterization order and the importance of locality-aware cooperative thread array scheduling for gaming applications.
Aditya Ukarande, Suryakant Patidar, Ram Rangan
ACM Trans. Archit. Code Optim.3
2021 PGZ: automatic zero-value code specialization
abstract
In prior work we proposed Zeroploit, a transform that duplicates code, specializes one path assuming certain key program operands, called versioning variables, are zero, and leaves the other path unspecialized. Dynamically, depending on the versioning variable’s value, either the specialized fast path or the default slow path will execute. We evaluated Zeroploit with hand-optimized codes in that work.
Mark Stephenson, Ram Rangan
CC2
2021 Cooperative Profile Guided Optimizations
abstract
Abstract Existing feedback‐driven optimization frameworks are not suitable for video games, which tend to push the limits of performance of gaming platforms and have real‐time constraints that preclude all but the simplest execution profiling. While Profile Guided Optimization (PGO) is a well‐established optimization approach, existing PGO techniques are ill‐suited for games for a number of reasons, particularly because heavyweight profiling makes interactive applications unresponsive. Adaptive optimization frameworks continually collect metrics that guide code specialization optimizations during program execution but have similarly high overheads. We emulate a system, which we callCooperative PGO, in which the gaming platform collectspiecemealprofiles by sampling in both time and space during actual gameplay across many users; stitches the piecemeal profiles together statistically; and creates policies to guide future gameplay. We introduce a three‐level hierarchical profiler that is well‐suited to graphics APIs, that commonly operates with no overhead and occasionally introduces an average overhead of less than 0.5% during periods of active profiling. This paper examines the practicality ofCooperative PGOusing three PGOs as case studies. A PGO that exploits likely zeros is particularly effective, achieving an average speedup of 5%, with a maximum speedup of 15%, over a highly‐tuned baseline.
Mark Stephenson, Ram Rangan, Stephen W. Keckler
Comput. Graph. Forum2
2020 Zeroploit: Exploiting Zero Valued Operands in Interactive Gaming Applications
abstract
In this article, we first characterize register operand value locality in shader programs of modern gaming applications and observe that there is a high likelihood of one of the register operands of several multiply, logical-and, and similar operations being zero, dynamically. We provide intuition, examples, and a quantitative characterization for how zeros originate dynamically in these programs. Next, we show that this dynamic behavior can be gainfully exploited with a profile-guided code optimization called Zeroploit that transforms targeted code regions into a zero-(value-)specialized fast path and a default slow path. The fast path benefits from zero-specialization in two ways, namely: (a) the backward slice of the other operand of a given multiply or logical-and can be skipped dynamically, provided the only use of that other operand is in the given instruction, and (b) the forward slice of instructions originating at the given instruction can be zero-specialized, potentially triggering further backward slice specializations from operations of that forward slice as well. Such specialization helps the fast path avoid redundant dynamic computations as well as memory fetches, while the fast-slow versioning transform helps preserve functional correctness. With an offline value profiler and manually optimized shader programs, we demonstrate that Zeroploit is able to achieve an average speedup of 35.8% for targeted shader programs, amounting to an average frame-rate speedup of 2.8% across a collection of modern gaming applications on an NVIDIA® GeForce RTX™ 2080 GPU.
Ram Rangan, Mark Stephenson, Aditya Ukarande, Shyam Murthy, Virat Agarwal, Marc Blackstein
ACM Trans. Archit. Code Optim.1
2013 Mesoscale performance simulation of multicore processor systems
Peter Altevogt, Tibor Kiss, Michael Kistler, Ram Rangan
Softw. Syst. Model.4
2010 Statistically regulating program behavior via mainstream computing
abstract
We introduce mainstream computing, a collaborative system that dynamically checks a program--via runtime assertion checks--to ensure that it is running according to expectation. Rather than enforcing strict, statically-defined assertions, our system allows users to run with a set of assertions that are statistically guaranteed to fail at a rate bounded by a user-defined probability, pfail. For example, a user can request a set of assertions that will fail at most 0.5% of the times the application is invoked. Users who believe their usage of an application is mainstream can use relatively large settings for pfail. Higher values of pfail provide stricter regulation of the application which likely enhances security, but will also inhibit some legitimate program behaviors; in contrast, program behavior is unregulated when pfail = 0, leaving the user vulnerable to attack. We show that our prototype is able to detect denial of service attacks, integer overflows, frees of uninitialized memory, boundary violations, and an injection attack. In addition we perform experiments with a mainstream computing system designed to protect against soft errors.
Mark Stephenson, Ram Rangan, Emmanuel Yashchin, Eric Van Hensbergen
CGO2
2009 Lightweight predication support for out of order processors
abstract
The benefits of Out of Order (OOO) processing are well known, as is the effectiveness of predicated execution for unpredictable control flow. However, as previous research has demonstrated, these techniques are at odds with one another. One common approach to reconciling their differences is to simplify the form of predication supported by the architecture. For instance, the only form of predication supported by modern OOO processors is a simple conditional move. We argue that it is the simplicity of conditional move that has allowed its widespread adoption, but we also show that this simplicity compromises its effectiveness as a compilation target. In this paper, we introduce a generalized form of hammock predication - called predicated mutually exclusive groups - that requires few modifications to an existing processor pipeline, yet presents the compiler with abundant predication opportunities. In comparison to non-predicated code running on an aggressively clocked baseline system, our technique achieves an 8% speedup averaged across three important benchmark suites.
Mark Stephenson, Lixin Zhang 0002, Ram Rangan
HPCA3
2008 Spice: speculative parallel iteration chunk execution
abstract
The recent trend in the processor industry of packing multiple processor cores in a chip has increased the importance of automatic techniques for extracting thread level parallelism. A promising approach for extracting thread level parallelism in general purpose applications is to apply memory alias or value speculation to break dependences amongst threads and executes them concurrently.
Easwaran Raman, Neil Vachharajani, Ram Rangan, David I. August
CGO3
2008 Performance scalability of decoupled software pipelining
abstract
Any successful solution to using multicore processors to scale general-purpose program performance will have to contend with rising intercore communication costs while exposing coarse-grained parallelism. Recently proposed pipelined multithreading (PMT) techniques have been demonstrated to have general-purpose applicability and are also able to effectively tolerate inter-core latencies through pipelined interthread communication. These desirable properties make PMT techniques strong candidates for program parallelization on current and future multicore processors and understanding their performance characteristics is critical to their deployment. To that end, this paper evaluates the performance scalability of a general-purpose PMT technique called decoupled software pipelining (DSWP) and presents a thorough analysis of the communication bottlenecks that must be overcome for optimal DSWP scalability.
Ram Rangan, Neil Vachharajani, Guilherme Ottoni, David I. August
ACM Trans. Archit. Code Optim.1
2007 Speculative Decoupled Software Pipelining
Neil Vachharajani, Ram Rangan, Easwaran Raman, Matthew J. Bridges, Guilherme Ottoni, David I. August
PACT2
2006 Support for High-Frequency Streaming in CMPs
abstract
As the industry moves toward larger-scale chip multiprocessors, the need to parallelize applications grows. High inter-thread communication delays, exacerbated by over-stressed high-latency memory subsystems and ever-increasing wire delays, require parallelization techniques to create partially or fully independent threads to improve performance. Unfortunately, developers and compilers alike often fail to find sufficient independent work of this kind. Recently proposed pipelined streaming techniques have shown significant promise for both manual and automatic parallelization. These techniques have wide-scale applicability because they embrace inter-thread dependences (albeit acyclic dependences) and tolerate long-latency communication of these dependences. This paper addresses the lack of architectural support for this type of concurrency, which has blocked its adoption and hindered related language and compiler research. We observe that both manual and automatic techniques create high-frequency streaming threads, with communication occurring every 5 to 20 instructions. Even while easily tolerating inter-thread transit delays, high-frequency communication makes thread performance very sensitive to intrathread delays from the repeated execution of the communication operations. Using this observation, we define the design-space and evaluate several mechanisms to find a better trade-off between performance and operating system, hardware, and design costs. From this, we find a light-weight streaming-aware enhancement to conventional memory subsystems that doubles the speed of these codes and is within 2% of the best-performing, but heavy-weight, hardware solution.
Ram Rangan, Neil Vachharajani, Adam Stoler, Guilherme Ottoni, David I. August, George Z. N. Cai
MICRO1
2005 SWIFT: Software Implemented Fault Tolerance
abstract
To improve performance and reduce power, processor designers employ advances that shrink feature sizes, lower voltage levels, reduce noise margins, and increase clock rates. However, these advances make processors more susceptible to transient faults that can affect correctness. While reliable systems typically employ hardware techniques to address soft-errors, software techniques can provide a lower-cost and more flexible alternative. This paper presents a novel, software-only, transient-fault-detection technique, called SWIFT. SWIFT efficiently manages redundancy by reclaiming unused instruction-level resources present during the execution of most programs. SWIFT also provides a high level of protection and performance with an enhanced control-flow checking mechanism. We evaluate an implementation of SWIFT on an Itanium 2 which demonstrates exceptional fault coverage with a reasonable performance cost. Compared to the best known single-threaded approach utilizing an ECC memory system, SWIFT demonstrates a 51% average speedup.
George A. Reis, Neil Vachharajani, Ram Rangan, David I. August
CGO4
2005 Computing Architectural Vulnerability Factors for Address-Based Structures
abstract
Processor designers require estimates of the architectural vulnerability factor (AVF) of on-chip structures to make accurate soft error rate estimates. AVF is the fraction of faults from alpha particle and neutron strikes that result in user-visible errors. This paper shows how to use a performance model to calculate the AVF of address-based structures, using a data cache, a data translation buffer, and a store buffer as examples. We describe how to perform a detailed breakdown of lifetime components (e.g., fill-to-read, read-to-evict) of bits in these structures into ACE (required for architecturally correct execution), un-ACE (unnecessary for ACE), and unknown components. This lifetime analysis produces best estimate AVFs for these three structures' data arrays of 6%, 36%, and 4%, respectively. We then present a new technique, hamming-distance-one analysis, and show that it predicts surprisingly low best estimate AVFs of 0.41%, 3%, and 7.7% for the structures' tag arrays. Finally, using our lifetime analysis framework, we show how two AVF reduction techniques - periodic flushing and incremental scrubbing - can reduce the AVF by converting ACE lifetime components into un-ACE without affecting performance significantly.
Arijit Biswas, Paul Racunas, Razvan Cheveresan, Joel S. Emer, Shubhendu S. Mukherjee, Ram Rangan
ISCA6
2005 Design and Evaluation of Hybrid Fault-Detection Systems
abstract
As chip densities and clock rates increase, processors are becoming more susceptible to transient faults that can affect program correctness. Up to now, system designers have primarily considered hardware-only and software-only fault-detection mechanisms to identify and mitigate the deleterious effects of transient faults. These two fault-detection systems, however, are extremes in the design space, representing sharp trade-offs between hardware cost, reliability, and performance. In this paper, we identify hybrid hardware/software fault-detection mechanisms as promising alternatives to hardware-only and software-only systems. These hybrid systems offer designers more options to fit their reliability needs within their hardware and performance budgets. We propose and evaluate CRAFT, a suite of three such hybrid techniques, to illustrate the potential of the hybrid approach. For fair, quantitative comparisons among hardware, software, and hybrid systems, we introduce a new metric, mean work to failure, which is able to compare systems for which machine instructions do not represent a constant unit of work. Additionally, we present a new simulation framework which rapidly assesses reliability and does not depend on manual identification of failure modes. Our evaluation illustrates that CRAFT, and hybrid techniques in general, offer attractive options in the fault-detection design space.
George A. Reis, Neil Vachharajani, Ram Rangan, David I. August, Shubhendu S. Mukherjee
ISCA4
2005 Automatic Thread Extraction with Decoupled Software Pipelining
abstract
Until recently, a steadily rising clock rate and other uniprocessor micro architectural improvements could be relied upon to consistently deliver increasing performance for a wide range of applications. Current difficulties in maintaining this trend have lead microprocessor manufacturers to add value by incorporating multiple processors on a chip. Unfortunately, since decades of compiler research have not succeeded in delivering automatic threading for prevalent code properties, this approach demonstrates no improvement for a large class of existing codes. To find useful work for chip multiprocessors, we propose an automatic approach to thread extraction, called decoupled software pipelining (DSWP). DSWP exploits the finegrained pipeline parallelism lurking in most applications to extract long-running, concurrently executing threads. Use of the nonspeculative and truly decoupled threads produced by DSWP can increase execution efficiency and provide significant latency tolerance, mitigating design complexity by reducing intercore communication and per-core resource requirements. Using our initial fully automatic compiler implementation and a validated processor model, we prove the concept by demonstrating significant gains for dual-core chip multiprocessor models running a variety of codes. We then explore simple opportunities missed by our initial compiler implementation which suggest a promising future for this approach.
Guilherme Ottoni, Ram Rangan, Adam Stoler, David I. August
MICRO2
2005 Software-controlled fault tolerance
abstract
Traditional fault-tolerance techniques typically utilize resources ineffectively because they cannot adapt to the changing reliability and performance demands of a system. This paper proposes software-controlled fault tolerance, a concept allowing designers and users to tailor their performance and reliability for each situation. Several software-controllable fault-detection techniques are then presented: SWIFT, a software-only technique, and CRAFT, a suite of hybrid hardware/software techniques. Finally, the paper introduces PROFiT, a technique which adjusts the level of protection and performance at fine granularities through software control. When coupled with software-controllable techniques like SWIFT and CRAFT, PROFiT offers attractive and novel reliability options.
George A. Reis, Neil Vachharajani, Ram Rangan, David I. August, Shubhendu S. Mukherjee
ACM Trans. Archit. Code Optim.4
2004 RIFLE: An Architectural Framework for User-Centric Information-Flow Security
abstract
Even as modern computing systems allow the manipulation and distribution of massive amounts of information, users of these systems are unable to manage the confidentiality of their data in a practical fashion. Conventional access control security mechanisms cannot prevent the illegitimate use of privileged data once access is granted. For example, information provided by a user during an online purchase may be covertly delivered to malicious third parties by an untrustworthy web browser. Existing information-flow security mechanisms do provide this assurance, but only for programmer-specified policies enforced during program development as a static analysis on special-purpose type-safe languages. Not only are these techniques not applicable to many commonly used programs, but they leave the user with no defense against malicious programmers or altered binaries. In this paper, we propose RIFLE, a runtime information-flow security system designed from the user's perspective. By addressing information-flow security using architectural support, RIFLE gives users a practical way to enforce their own information-flow security policy on all programs. We prove that, contrary to statements in the literature, run-time systems like RIFLE are no less secure than existing language-based techniques. Using a model of the architectural framework and a binary translator, we demonstrate RIFLE's correctness and illustrate that the performance cost is reasonable.
Neil Vachharajani, Matthew J. Bridges, Ram Rangan, Guilherme Ottoni, Jason A. Blome, George A. Reis, Manish Vachharajani, David I. August
MICRO4