VLDB 2026 Research / reviewers in the wild / expert
Sanjay J. Patel
dblp:05/6203
· DBLP profile ↗
42ranked-venue papers
5as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 5 first-authorSoftware engineering, systems software and programming languages · 14 · 1 first-authorSecurity and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
31 papers |
Processor architecture and microarchitecture · 41% Hardware reliability and fault tolerance · 9% GPUs and heterogeneous computing · 8% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 100% |
Topics — the 30 heaviest of 63, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
memory latency tolerance |
0.3 | 2 | 2013 | Hybrid latency tolerance for robust energy-efficiency on 1000-core data parallel processors · HPCA 2013 OUTRIDER: efficient memory latency tolerance with decoupled strands · ISCA 2011 |
Parallel and multicore computing
parallel programming models |
0.2 | 3 | 2010 | Rigel: an architecture and scalable programming interface for a 1000-core accelerator · ISCA 2009 Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 4 | 2006 | Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006 Dynamic Optimization of Micro-Operations · HPCA 2003 Increasing the size of atomic instruction blocks using control flow assertions · MICRO 2000 |
Processor architecture and microarchitecture
dynamic optimization |
0.1 | 4 | 2005 | Continuous Optimization · ISCA 2005 rePLay: A Hardware Framework for Dynamic Optimization · IEEE Trans. Computers 2001 Performance characterization of a hardware mechanism for dynamic optimization · MICRO 2001 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 2 | 2007 | Examining ACE analysis reliability estimates using fault-injection · ISCA 2007 ReStore: Symptom-Based Soft Error Detection in Microprocessors · IEEE Trans. Dependable Secur. Comput. 2006 |
Processor architecture and microarchitecture
throughput processor |
0.1 | 1 | 2011 | OUTRIDER: efficient memory latency tolerance with decoupled strands · ISCA 2011 |
Processor architecture and microarchitecture
instruction fetch |
0.1 | 5 | 2002 | Instruction fetch deferral using static slack · MICRO 2002 Evaluation of Design Options for the Trace Cache Fetch Mechanism · IEEE Trans. Computers 1999 Putting the Fill Unit to Work: Dynamic Optimizations for Trace Cache Microprocessors · MICRO 1998 |
Processor architecture and microarchitecture
pipelining |
0.1 | 2 | 2006 | Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006 Reducing the Scheduling Critical Cycle Using Wakeup Prediction · HPCA 2004 |
Performance modeling and evaluation
analytical modeling |
0.1 | 1 | 2010 | An adaptive performance modeling tool for GPU architectures · PPoPP 2010 |
Memory systems
cache coherence |
0.1 | 1 | 2010 | Cohesion: a hybrid memory model for accelerators · ISCA 2010 |
Energy-efficient computing
circuit-architecture co-optimization |
0.1 | 1 | 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis · ISCA 2010 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Performance modeling and evaluation › performance prediction
GPU performance prediction |
0.1 | 1 | 2010 | An adaptive performance modeling tool for GPU architectures · PPoPP 2010 |
GPUs and heterogeneous computing
heterogeneous parallel programming |
0.1 | 1 | 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systems · ASPLOS 2010 |
Electronic design automation › design space exploration
microarchitecture design space exploration |
0.1 | 1 | 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis · ISCA 2010 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 1 | 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis · ISCA 2010 |
Energy-efficient computing
voltage scaling |
0.1 | 1 | 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysis · ISCA 2010 |
Hardware accelerators and domain-specific architectures
many-core accelerator |
0.1 | 1 | 2009 | Rigel: an architecture and scalable programming interface for a 1000-core accelerator · ISCA 2009 |
Hardware accelerators and domain-specific architectures › accelerator architecture
programmable accelerator |
0.1 | 1 | 2009 | Rigel: an architecture and scalable programming interface for a 1000-core accelerator · ISCA 2009 |
Processor architecture and microarchitecture › instruction fetch
trace cache |
0.1 | 5 | 2002 | Evaluation of Design Options for the Trace Cache Fetch Mechanism · IEEE Trans. Computers 1999 Putting the Fill Unit to Work: Dynamic Optimizations for Trace Cache Microprocessors · MICRO 1998 Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 4 | 2000 | Increasing the size of atomic instruction blocks using control flow assertions · MICRO 2000 Evaluation of Design Options for the Trace Cache Fetch Mechanism · IEEE Trans. Computers 1999 Improving Trace Cache Effectiveness with Branch Promotion and Trace Packing · ISCA 1998 |
Processor architecture and microarchitecture
instruction scheduling |
0.1 | 2 | 2004 | Reducing the Scheduling Critical Cycle Using Wakeup Prediction · HPCA 2004 Instruction fetch deferral using static slack · MICRO 2002 |
Performance modeling and evaluation
workload characterization |
0.1 | 3 | 2008 | Characterization of essential dynamic instructions · SIGMETRICS 2003 Tradeoffs in designing accelerator architectures for visual computing · MICRO 2008 An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors Work · ISCA 1998 |
Integrated circuit design › digital circuit design › arithmetic circuit design
floating-point unit |
0.1 | 1 | 2007 | The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics Acceleration · MICRO 2007 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Emerging computing paradigms › approximate computing
precision reduction |
0.1 | 1 | 2007 | The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics Acceleration · MICRO 2007 |
Hardware reliability and fault tolerance
reliability analysis |
0.1 | 1 | 2007 | Examining ACE analysis reliability estimates using fault-injection · ISCA 2007 |
Hardware reliability and fault tolerance › soft errors
transient fault |
0.1 | 1 | 2007 | Examining ACE analysis reliability estimates using fault-injection · ISCA 2007 |
Processor architecture and microarchitecture
latency hiding |
0.1 | 1 | 2006 | Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006 |
Memory systems › memory access latency
load latency |
0.1 | 1 | 2006 | Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.3performance and physical design models · 0.2design space exploration · 0.2marginal cost analysis · 0.1fine-grained temporal coherence reassignment · 0.1data transfer management · 0.1architecture-circuit optimization framework · 0.1analytical modeling · 0.1level-of-detail · 0.1LCP solver · 0.1instruction marking · 0.1store forwarding · 0.1redundant load elimination · 0.1reassociation · 0.1constant propagation · 0.1frame formation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Hybrid latency tolerance for robust energy-efficiency on 1000-core data parallel processorsabstractCurrently, GPUs and data parallel processors leverage latency tolerance techniques such as multithreading and prefetching to maximize performance per Watt. However, choosing a technique that provides energy-efficiency on a wide variety of workloads is difficult, as the type of latency to tolerate, required hardware complexity, and energy consumption is directly related to application behavior. After qualitatively evaluating five commonly used latency tolerance techniques, we develop a hybrid technique utilizing multithreading and decoupled execution to maximize performance while minimizing hardware complexity and energy consumption across a wide variety of workloads. We compare our hybrid technique with the five commonly used techniques on a 1024-core data parallel processor by performing a comprehensive design space exploration, leveraging detailed performance and physical design models. By intelligently leveraging both decoupled execution and multithreading, our hybrid latency tolerance technique is able to improve energy-efficiency by 28% to 89% over any single technique on data parallel benchmarks. Compared to other combinations of latency tolerance techniques, we find that our hybrid latency tolerance technique provides the highest energy-efficiency by over 26%. Neal Clayton Crago, Omid Azizi, Steven S. Lumetta, Sanjay J. Patel |
HPCA | 4 |
| 2011 | Decoupled Architectures as a Low-Complexity Alternative to Out-of-order ExecutionabstractIn this paper we present OUTRIDERHP, a novel implementation of a decoupled architecture that approaches the performance of contemporary out-of-order processors on parallel benchmarks while maintaining low hardware complexity. OUTRIDERHP leverages the compiler to separate a single thread of execution into memory-accessing and memory-consuming streams that can be executed concurrently, which we call strands. We identify loss-of-decoupling events which cripple performance on traditional decoupled architectures, and design OUTRIDERHP to enable extraction of multiple strands and control speculation which provide superior memory and functional unit latency tolerance. OUTRIDERHP outperforms a baseline in-order architecture by 26-220% and Decoupled Access/Execute by 7-172% when executing parallel benchmarks on an 8-core CMP configuration. OUTRIDERHP performs within 15% of higher-complexity out-of-order cores despite not utilizing large physical register files, dynamic scheduling, and register renaming hardware. Neal Clayton Crago, Sanjay J. Patel |
PACT | 2 |
| 2011 | Accelerating aerial image simulation with GPUabstractAerial image simulation is a fundamental problem for modern VLSI design. It requires a huge amount of numerical computation. The recent advancement of general purpose GPU computing provides an excellent opportunity to parallelize the aerial image simulation and achieve great speedup. In this paper, we present and discuss two GPU-based aerial image simulation algorithms. We show through experiments that the fastest algorithm we propose can achieve 50X to 60X speedup over the CPU based serial algorithm. The error of our approach is shown to be insignificant. Hongbo Zhang 0001, Tan Yan, Martin D. F. Wong, Sanjay J. Patel |
ICCAD | 4 |
| 2011 | OUTRIDER: efficient memory latency tolerance with decoupled strandsabstractWe present OUTRIDER, an architecture for throughput-oriented processors that provides memory latency tolerance to improve performance on highly threaded workloads. OUTRIDER enables a single thread of execution to be presented to the architecture as multiple decoupled instruction streams that separate memory-accessing and memory-consuming instructions. The key insight is that by decoupling the instruction streams, the processor pipeline can tolerate memory latency in a way similar to out-of-order designs while relying on a low-complexity in-order micro-architecture. Moreover, instead of adding more threads as is done in modern GPUs, OUTRIDER can tolerate memory latency with fewer threads and reduced contention for resources shared amongst threads. We demonstrate that OUTRIDER can outperform single threaded cores by 23-131% and a 4-way simultaneous multithreaded core by up to 87% on data parallel applications in a 1024-core system. Moreover, OUTRIDER achieves these performance gains without incurring the overhead of additional hardware thread contexts, which results in improved area efficiency compared to a multithreaded core. Neal Clayton Crago, Sanjay J. Patel |
ISCA | 2 |
| 2010 | WAYPOINT: scaling coherence to thousand-core architecturesabstractIn this paper, we evaluate a set of coherence architectures in the context of a 1024-core chip multiprocessor (CMP) tailored to throughput-oriented parallel workloads. Based on our analysis, we develop and evaluate two techniques for scaling coherence to thousand-core CMPs. We find that a broadcast-based probe filtering scheme provides reasonable performance up to 128 cores for some benchmarks, but is not generally scalable. We propose a broadcast-collective network for accelerating probe filter misses, which extends scalability but falls short of supporting 1024 cores. We find that a sparse directory with an invalidate-on-evict policy can work well for many throughput-oriented workloads. However, the on-die structures required to achieve good performance carry a large performance and power overhead. To achieve thousand-core scalability with smaller and less associative sparse directories, we introduce WayPoint, a mechanism that increases directory associativity and capacity dynamically. Using less than 3% of total die area, Way-Point achieves performance within 4% of an infinitely large on-die directory. John H. Kelm, Matthew R. Johnson 0003, Steven S. Lumetta, Sanjay J. Patel |
PACT | 4 |
| 2010 | An asymmetric distributed shared memory model for heterogeneous parallel systemsabstractHeterogeneous computing combines general purpose CPUs with accelerators to efficiently execute both sequential control-intensive and data-parallel phases of applications. Existing programming models for heterogeneous computing rely on programmers to explicitly manage data transfers between the CPU system memory and accelerator memory. Isaac Gelado, Javier Cabezas, Nacho Navarro, John E. Stone, Sanjay J. Patel, Wen-Mei W. Hwu |
ASPLOS | 5 |
| 2010 | An integrated framework for joint design space exploration of microarchitecture and circuitsabstractThe design of a digital system for energy efficiency often requires the analysis of circuit tradeoffs in addition to architectural tradeoffs. To assist with this analysis, we present a framework for performing joint exploration of both the architectural and circuit design spaces. In our approach, we use statistical inference techniques to create a model of a large micro-architectural design space from a small number of simulation samples. We then characterize the design tradeoffs of each of the underlying circuits and integrate these with the higher level architectural models to define the joint circuit-architecture design space. We use posynomial forms for all our models, enabling the use of convex optimization tools to efficiently search the joint design space. As an example, we apply this methodology to explore the power-performance tradeoffs in a dual-issue superscalar out-of-order processor, showing how the framework can be used to determine the optimal set of design parameters for energy efficiency. Compared to current architectural tools that use fixed circuit costs, joint optimization can reduce energy by up to 30% by considering circuit tradeoff characteristics. Omid Azizi, Aqeel Mahesri, John P. Stevenson, Sanjay J. Patel, Mark Horowitz |
DATE | 4 |
| 2010 | GoldMine: Automatic assertion generation using data mining and static analysisabstractWe present GOLDMINE, a methodology for generating assertions automatically. Our method involves a combination of data mining and static analysis of the Register Transfer Level (RTL) design. We present results of using GoldMine for assertion generation of the RTL of a 1000-core processor design that is still in an evolving stage. Our results show that GoldMine can generate complex, high coverage assertions in RTL, thereby minimizing human effort in this process. Shobha Vasudevan, David Sheridan, Sanjay J. Patel, David Tcheng, William Tuohy, Daniel R. Johnson |
DATE | 3 |
| 2010 | Energy-performance tradeoffs in processor architecture and circuit design: a marginal cost analysisabstractPower consumption has become a major constraint in the design of processors today. To optimize a processor for energy-efficiency requires an examination of energy-performance trade-offs in all aspects of the processor design space, including both architectural and circuit design choices. In this paper, we apply an integrated architecture-circuit optimization framework to map out energy-performance trade-offs of several different high-level processor architectures. We show how the joint architecture-circuit space provides a trade-off range of approximately 6.5x in performance for 4x energy, and we identify the optimal architectures for different design objectives. We then show that many of the designs in this space come at very high marginal costs. Our results show that, for a large range of design objectives, voltage scaling is effective in efficiently trading off performance and energy, and that the choice of optimal architecture and circuits does not change much during voltage scaling. Finally, we show that with only two designs--a dual-issue in-order design and a dual-issue out-of-order design, both properly optimized-a large part of the energy-performance trade-off space can be covered within 3% of the optimal energy-efficiency. Omid Azizi, Aqeel Mahesri, Benjamin C. Lee, Sanjay J. Patel, Mark Horowitz |
ISCA | 4 |
| 2010 | Cohesion: a hybrid memory model for acceleratorsabstractTwo broad classes of memory models are available today: models with hardware cache coherence, used in conventional chip multiprocessors, and models that rely upon software to manage coherence, found in compute accelerators. In some systems, both types of models are supported using disjoint address spaces and/or physical memories. In this paper we present Cohesion, a hybrid memory model that enables fine-grained temporal reassignment of data between hardware-managed and software-managed coherence domains, allowing a system to support both. Cohesion can be used to dynamically adapt to the sharing needs of both applications and runtimes. Cohesion requires neither copy operations nor multiple address spaces. John H. Kelm, Daniel R. Johnson, William Tuohy, Steven S. Lumetta, Sanjay J. Patel |
ISCA | 5 |
| 2010 | An adaptive performance modeling tool for GPU architecturesabstractThis paper presents an analytical model to predict the performance of Sara S. Baghsorkhi, Matthieu Delahaye, Sanjay J. Patel, William Gropp, Wen-Mei W. Hwu |
PPoPP | 3 |
| 2009 | A Task-Centric Memory Model for Scalable Accelerator ArchitecturesabstractThis paper presents a task-centric memory model for 1000-core compute accelerators. Visual computing applications are emerging as an important class of workloads that can exploit 1000-core processors. In these workloads, we observe data sharing and communication patterns that can be leveraged in the design of memory systems for future 1000-core processors. Based on these insights, we propose a memory model that uses a software protocol, working in collaboration with hardware caches, to maintain a coherent, single-address space view of memory without the need for hardware coherence support. We evaluate the task-centric memory model in simulation on a 1024-core MIMD accelerator we are developing that, with the help of a runtime system, implements the proposed memory model. We evaluate coherence management policies related to the task-centric memory model and show that the overhead of maintaining a coherent view of memory in software can be minimal. We further show that, while software management may constrain speculative hardware prefetching into local caches, a common optimization, it does not constrain the more relevant use case of off-chip prefetching from DRAM into shared caches. John H. Kelm, Daniel R. Johnson, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel |
PACT | 5 |
| 2009 | Depth image-based rendering with low resolution depthabstractThis paper proposes a new approach for depth image-based rendering (DIBR) with low resolution depth using the 3D propagation algorithm. Our novel depth edge enhancement method efficiently corrects and sharpens the depth edges in the propagated depth image using available high resolution color information. Experimental results show that only with 4% depth information kept for low resolution depth image, the proposed method can provide comparable rendering quality to that of the high resolution case. Furthermore, the proposed work is developed to match the fine-grain parallelism of general-purpose graphics processing units (GPGPUs) and hence can be accelerated to nearly real-time operations in low cost DIBR systems. Minh N. Do, Sanjay J. Patel |
ICIP | 3 |
| 2009 | Rigel: an architecture and scalable programming interface for a 1000-core acceleratorabstractThis paper considers Rigel, a programmable accelerator architecture for a broad class of data- and task-parallel computation. Rigel comprises 1000+ hierarchically-organized cores that use a fine-grained, dynamically scheduled single-program, multiple-data (SPMD) execution model. Rigel's low-level programming interface adopts a single global address space model where parallel work is expressed in a task-centric, bulk-synchronized manner using minimal hardware support. Compared to existing accelerators, which contain domain-specific hardware, specialized memories, and/or restrictive programming models, Rigel is more flexible and provides a straightforward target for a broader set of applications. John H. Kelm, Daniel R. Johnson, Matthew R. Johnson 0003, Neal Clayton Crago, William Tuohy, Aqeel Mahesri, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel |
ISCA | 9 |
| 2009 | Fool me twice: Exploring and exploiting error tolerance in physics-based animationabstractThe error tolerance of human perception offers a range of opportunities to trade numerical accuracy for performance in physics-based simulation. However, most prior work on perceptual error tolerance either focus exclusively on understanding the tolerance of the human visual system or burden the application developer with case-specific implementations such as Level-of-Detail (LOD) techniques. In this article, based on a detailed set of perceptual metrics, we propose a methodology to identify the maximum error tolerance of physics simulation. Then, we apply this methodology in the evaluation of four case studies. First, we utilize the methodology in the tuning of the simulation timestep. The second study deals with tuning the iteration count for the LCP solver. Then, we evaluate the perceptual quality of Fast Estimation with Error Control (FEEC) [Yeh et al. 2006]. Finally, we explore the hardware optimization technique of precision reduction. Thomas Y. Yeh, Glenn Reinman, Sanjay J. Patel, Petros Faloutsos |
ACM Trans. Graph. | 3 |
| 2008 | Tradeoffs in designing accelerator architectures for visual computingabstractVisualization, interaction, and simulation (VIS) constitute a class of applications that is growing in importance. This class includes applications such as graphics rendering, video encoding, simulation, and computer vision. These applications are ideally suited for accelerators because of their parallelizability and demand for high throughput. We compile a benchmark suite, VIS- Bench, to serve as a proxy for this application class. We use VISBench to examine some important high level decisions for an accelerator architecture. We propose a highly parallel base architecture. We examine the need for synchronization and data communication. We also examine GPU-style SIMD execution and find that a MIMD architecture usually performs better. Given these high level choices, we use VISBench to explore the microarchitectural design space. We analyze area versus performance tradeoffs in designing individual cores and the memory hierarchy. We find that a design made of small, simple cores achieves much higher throughput than a general purpose uniprocessor. Further, we find that a limited amount of support for ILP within each core aids overall performance. We find that fine-grained multithreading improves performance, but only up to a point. We find that word-level (SSE-style) SIMD provides a poor performance to area ratio. Finally, we find that sufficient memory and cache bandwidth is essential to performance. Aqeel Mahesri, Daniel R. Johnson, Neal Clayton Crago, Sanjay J. Patel |
MICRO | 4 |
| 2007 | Implicitly Parallel Programming Models for Thousand-Core MicroprocessorsabstractThis paper argues for an implicitly parallel programming model for many-core microprocessors, and provides initial technical approaches towards this goal. In an implicitly parallel programming model, programmers maximize algorithm-level parallelism, express their parallel algorithms by asserting high-level properties on top of a traditional sequential programming language, and rely on parallelizing compilers and hardware support to perform parallel execution under the hood. In such a model, compilers and related tools require much more advanced program analysis capabilities and programmer assertions than what are currently available so that a comprehensive understanding of the input program's concurrency can be derived. Such an understanding is then used to drive automatic or interactive parallel code generation tools for a diverse set of parallel hardware organizations. The chip-level architecture and hardware should maintain parallel execution state in such a way that a strictly sequential execution state can always be derived for the purpose of verifying and debugging the program. We argue that implicitly parallel programming models are critical for addressing the software development crises and software scalability challenges for many-core microprocessors. Wen-Mei W. Hwu, Shane Ryoo, Sain-Zee Ueng, John H. Kelm, Isaac Gelado, Sam S. Stone, Robert E. Kidd, Sara S. Baghsorkhi, Aqeel Mahesri, Stephanie C. Tsao, Nacho Navarro, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel |
DAC | 14 |
| 2007 | Examining ACE analysis reliability estimates using fault-injectionabstractACE analysis is a technique to provide an early reliability estimate for microprocessors. ACE analysis couples data from abstract performance models with low level design details to identify and rule out transient faults that will not cause incorrect execution. While many transient faults are analyzable in ACE analysis frameworks, some are not. As a result, ACE analysis is conservative and provides a lower bound for the reliability of a processor design. Bounding the reliability of a design is useful since it can guarantee that the given design will meet reliability goals. Nicholas J. Wang, Aqeel Mahesri, Sanjay J. Patel |
ISCA | 3 |
| 2007 | ParallAX: an architecture for real-time physicsabstractFuture interactive entertainment applications will featurethe physical simulation of thousands of interacting objectsusing explosions, breakable objects, and cloth effects. Whilethese applications require a tremendous amount of performanceto satisfy the minimum frame rate of 30 FPS, there is a dramatic amount of parallelism in future physics workloads.How will future physics architectures leverage parallelismto achieve the real-time constraint?. Thomas Y. Yeh, Petros Faloutsos, Sanjay J. Patel, Glenn Reinman |
ISCA | 3 |
| 2007 | The Art of Deception: Adaptive Precision Reduction for Area Efficient Physics AccelerationabstractPhysics-based animation has enormous potential to improve the realism of interactive entertainment through dynamic, immersive content creation. Despite the massively parallel nature of physics simulation, fully exploiting this parallelism to reach interactive frame rates will require significant area to place the large number of cores. Fortunately, interactive entertainment requires believability rather than accuracy. Recent work shows that real-time physics has a remarkable tolerance for reduced precision of the significant in floating-point (FP) operations. In this paper, we describe an architecture with a hierarchical floating-point unit (FPU) that leverages dynamic precision reduction to enable efficient FPU sharing among multiple cores. This sharing reduces the area required by these cores, thereby allowing more cores to be packed into a given area and exploiting more parallelism. Thomas Y. Yeh, Petros Faloutsos, Milos D. Ercegovac, Sanjay J. Patel, Glenn Reinman |
MICRO | 4 |
| 2006 | Beating In-Order Stalls with "Flea-Flicker" Two-Pass PipeliningabstractWhile compilers have generally proven adept at planning useful static instruction-level parallelism for in-order microarchitectures, the efficient accommodation of unanticipateable latencies, like those of load instructions, remains a vexing problem. Traditional out-of-order execution hides some of these latencies, but repeats scheduling work already done by the compiler and adds additional pipeline overhead. Other techniques, such as prefetching and multithreading, can hide some anticipateable, long-latency misses, but not the shorter, more diffuse stalls due to difficult-to-anticipate, first or second-level misses. Our work proposes a microarchitectural technique, two-pass pipelining, whereby the program executes on two in-order back-end pipelines coupled by a queue. The "advance" pipeline often defers instructions dispatching with unready operands rather than stalling. The "backup" pipeline allows concurrent resolution of instructions deferred by the first pipeline allowing overlapping of useful "advanced" execution with miss resolution. An accompanying compiler technique and instruction marking further enhance the handling of miss latencies. Applying our technique to an Itanium 2-like design achieves a speedup of 1.38x in mcf, the most memory-intensive SPECint2000 benchmark, and an average of 1.12 x across other selected benchmarks, yielding between 32 percent and 67 percent of an idealized out-of-order design's speedup at a much lower design cost and complexity. Ronald D. Barnes, John W. Sias, Erik M. Nystrom, Sanjay J. Patel, Nacho Navarro, Wen-Mei W. Hwu |
IEEE Trans. Computers | 4 |
| 2006 | ReStore: Symptom-Based Soft Error Detection in MicroprocessorsabstractDevice scaling and large-scale integration have led to growing concerns about soft errors in microprocessors. To date, in all but the most demanding applications, implementing parity and ECC for caches and other large, regular SRAM structures have been sufficient to stem the growing soft error tide. This will not be the case for long and questions remain as to the best way to detect and recover from soft errors in the remainder of the processor - in particular, the less structured execution core. In this work, we propose the ReStore architecture, which leverages existing performance enhancing checkpointing hardware to recover from soft error events in a low cost fashion. Error detection in the ReStore architecture is novel: symptoms that hint at the presence of soft errors trigger restoration of a previous checkpoint. Example symptoms include exceptions, control flow misspeculations, and cache or translation look-aside buffer misses. Compared to conventional soft error detection via full replication, the ReStore framework incurs little overhead, but sacrifices some amount of error coverage. These attributes make it an ideal means to provide very cost effective error coverage for processor applications that can tolerate a nonzero, but small, soft error failure rate. Our evaluation of an example ReStore implementation exhibits a 2times increase in MTBF (mean time between failures) over a standard pipeline with minimal hardware and performance overheads. The MTBF increases by 20times if ReStore is coupled with protection for certain particularly vulnerable pipeline structures Nicholas J. Wang, Sanjay J. Patel |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2006 | Sequential Element Design With Built-In Soft Error ResilienceabstractThis paper presents a built-in soft error resilience (BISER) technique for correcting radiation-induced soft errors in latches and flip-flops. The presented error-correcting latch and flip-flop designs are power efficient, introduce minimal speed penalty, and employ reuse of on-chip scan design-for-testability and design-for-debug resources to minimize area overheads. Circuit simulations using a sub-90-nm technology show that the presented designs achieve more than a 20-fold reduction in cell-level soft error rate (SER). Fault injection experiments conducted on a microprocessor model further demonstrate that chip-level SER improvement is tunable by selective placement of the presented error-correcting designs. When coupled with error correction code to protect in-pipeline memories, the BISER flip-flop design improves chip-level SER by 10 times over an unprotected pipeline with the flip-flops contributing an extra 7-10.5% in power. When only soft errors in flips-flops are considered, the BISER technique improves chip-level SER by 10 times with an increased power of 10.3%. The error correction mechanism is configurable (i.e., can be turned on or off) which enables the use of the presented techniques for designs that can target multiple applications with a wide range of reliability requirements Ming Zhang 0017, Subhasish Mitra, Norbert Seifert, Nicholas J. Wang, Kee Sup Kim, Naresh R. Shanbhag, Sanjay J. Patel |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2005 | ReStore: Symptom Based Soft Error Detection in MicroprocessorsabstractDevice scaling and large scale integration have led to growing concerns about soft errors in microprocessors. To date, in all but the most demanding applications, implementing parity and ECC for caches and other large, regular SRAM structures have been sufficient to stem the growing soft error tide. This will not be the case for long, and questions remain as to the best way to detect and recover from soft errors in the remainder of the processor - in particular, the less structured execution core. In this work, we propose the ReStore architecture, which leverages existing performance enhancing checkpointing hardware to recover from soft error events in a low cost fashion. Error detection in the ReStore architecture is novel: symptoms that hint at the presence of soft errors trigger restoration of a previous checkpoint. Example symptoms include exceptions, control flow mis-speculations, and cache or translation look-aside buffer misses. Compared to conventional soft error detection via full replication, the ReStore framework incurs little overhead, but sacrifices some amount of error coverage. These attributes make it an ideal means to provide very cost effective error coverage for processor applications that can tolerate a nonzero, but small, soft error failure rate. Our evaluation of an example ReStore implementation exhibits a 2x increase in MTBE (mean time between failures) over a standard pipeline with minimal hardware and performance overheads. The MTBF increases by 7x if ReStore is coupled with parity protection for certain pipeline structures. Nicholas J. Wang, Sanjay J. Patel |
DSN | 2 |
| 2005 | The Future of Computer Architecture Research: An Industrial PerspectiveabstractSummary form only given. This panel focuses on the industrial vision for the future of computer architecture research. The field of computer architecture is in critical need for focus, perhaps now more than ever. The underlying technology is presenting significant roadblocks for next generation designs. The applications that the world wants seem to be out of synch with mainstream research. The potential impact of our work seems to be uncertain. Despite this, there are some excellent opportunities out there. We have invited several distinguished technical leaders from key hardware and software companies in the computing industry to articulate their vision for the future of computer architecture research. Wen-Mei W. Hwu, Sanjay J. Patel |
HPCA | 2 |
| 2005 | Continuous OptimizationabstractThis paper presents a hardware-based dynamic optimizer that continuously optimizes an application's instruction stream. In continuous optimization, dataflow optimizations are performed using simple, table-based hardware placed in the rename stage of the processor pipeline. The continuous optimizer reduces dataflow height by performing constant propagation, reassociation, redundant load elimination, store forwarding, and silent store removal. To enhance the impact of the optimizations, the optimizer integrates values generated by the execution units back into the optimization process. Continuous optimization allows instructions with input values known at optimization time to be executed in the optimizer, leaving less work for the out-of-order portion of the pipeline. Continuous optimization can detect branch mispredictions earlier and thus reduce the misprediction penalty. In this paper, we present a detailed description of a hardware optimizer and evaluate it in the context of a contemporary microarchitecture running current workloads. Our analysis of SPECint, SPECfp, and mediabench workloads reveals that a hardware optimizer can directly execute 33% of instructions, resolve 29% of mispredicted branches, and generate addresses for 76% of memory operations. These positive effects combine to provide speed ups in the range 0.99 to 1.27. Brian Fahs, Todd M. Rafacz, Sanjay J. Patel, Steven S. Lumetta |
ISCA | 3 |
| 2004 | Characterizing the Effects of Transient Faults on a High-Performance Processor PipelineabstractThe progression of implementation technologies into the sub-100 nanometer lithographies renew the importance of understanding and protecting against single-event upsets in digital systems. In this work, the effects of transient faults on high performance microprocessors is explored. To perform a thorough exploration, a highly detailed register transfer level model of a deeply pipelined, out-of-order microprocessor was created. Using fault injection, we determined that fewer than 15% of single bit corruptions in processor state result in software visible errors. These failures were analyzed to identify the most vulnerable portions of the processor, which were then protected using simple low-overhead techniques. This resulted in a 75% reduction in failures. Building upon the failure modes seen in the microarchitecture, fault injections into software were performed to investigate the level of masking that the software layer provides. Together, the baseline microarchitectural substrate and software mask more than 9 out of 10 transient faults from affecting correct program execution. Nicholas J. Wang, Justin Quek, Todd M. Rafacz, Sanjay J. Patel |
DSN | 4 |
| 2004 | Reducing the Scheduling Critical Cycle Using Wakeup PredictionabstractFor highest performance, a modern microprocessor must be able to determine if an instruction is ready in the same cycle in which it is to be selected for execution. This creates a cycle of logic involving wakeup and select. However, the time a static instruction spends waiting for wakeup shows little dynamic variance. This idea is used to build a machine where wakeup times are predicted, and instructions executed too early are replayed. This form of self-scheduling reduces the critical cycle by eliminating the wakeup logic at the expense of additional replays. However, replays and other pipeline effects affect the cost of misprediction. To solve this, an allowance is added to the predicted wakeup time to decrease the probability of a replay. This allowance may be associated with individual instructions or the global state, and is dynamically adjusted by a gradient-descent minimum-searching technique. When processor load is low, prediction may be more aggressive — increasing the chance of replays, but increasing performance, so the aggressiveness of the predictor is dynamically adjusted using processor load as a feedback parameter. Todd E. Ehrhart, Sanjay J. Patel |
HPCA | 2 |
| 2003 | Improving Quasi-Dynamic Schedules through Region SlipabstractModern processors perform dynamic scheduling to achieve better utilization of execution resources. A schedule created at run-time is often better than one created at compile-time as it can dynamically adapt to specific events encountered at execution-time. In this paper, we examine some fundamental impediments to effective static scheduling. More specifically, we examine the question of why schedules generated quasi-dynamically by a low-level runtime optimizer and executed on a statically scheduled machine perform worse than using a dynamically-scheduled approach. We observe that such schedules suffer because of region boundaries and a skewed distribution of parallelism towards the beginning of a region. To overcome these limitations, we investigate a new concept, region slip, in which the schedules of different statically-scheduled regions can be interleaved in the processor issue queue to reduce the region boundary effects that cause empty issue slots. Francesco Spadini, Brian Fahs, Sanjay J. Patel, Steven S. Lumetta |
CGO | 3 |
| 2003 | Dynamic Optimization of Micro-OperationsabstractInherent within complex instruction set architectures such as /spl times/86 are inefficiencies that do not exist in a simpler ISA. Modern /spl times/86 implementations decode instructions into one or more micro-operations in order to deal with the complexity of the ISA. Since these micro-operations are not visible to the compiler the stream of micro-operations can contain redundancies even in statically optimized /spl times/86 code. Within a processor implementation, however barriers at the ISA level do not apply, and these redundancies can be removed by optimizing the micro-operation stream. In this paper we explore the opportunities to optimize code at the micro-operation granularity. We execute these micro-operation optimizations using the rePLay Framework as a microarchitectural substrate. Using a simple set of seven optimizations, including two that aggressively and speculatively attempt to remove redundant load instructions, we examine the effects of dynamic optimization of micro-operations using a trace-driven simulation environment. Simulation reveals that across a sampling of SPECint 2000 and real /spl times/86 applications, rePLay is able to reduce micro-operation count by 21% and, in particular load micro-operation count by 22%. These reductions correspond to a boost in observed instruction-level parallelism on an 8-wide optimizing rePLay processor by 17% over a non-optimizing configuration. Brian Slechta, David Crowe, Brian Fahs, Michael Fertig, Gregory A. Muthler, Justin Quek, Francesco Spadini, Sanjay J. Patel, Steven S. Lumetta |
HPCA | 8 |
| 2003 | Beating in-order stalls with "flea-flicker" two-pass pipeliningabstractAccommodating the uncertain latency of load instructions is one of the most vexing problems in in-order microarchitecture design and compiler development. Compilers can generate schedules with a high degree of instruction-level parallelism but cannot effectively accommodate unanticipated latencies; incorporating traditional out-of-order execution into the microarchitecture hides some of this latency but redundantly performs work done by the compiler and adds additional pipeline stages. Although effective techniques, such as prefetching and threading, have been proposed to deal with anticipable, long latency misses, the shorter, more diffuse stalls due to difficult-to-anticipate, first- or second-level misses are less easily hidden on in-order architectures. This paper addresses this problem by proposing a microarchitectural technique, referred to as two-pass pipelining, wherein the program executes on two in-order back-end pipelines coupled by a queue. The "advance" pipeline executes instructions greedily, without stalling on unanticipated latency dependences (executing independent instructions while otherwise blocking instructions are deferred). The "backup" pipeline allows concurrent resolution of instructions that were deferred in the other pipeline, resulting in the absorption of shorter misses and the overlap of longer ones. This paper argues that this design is both achievable and a good use of transistor resources and shows results indicating that it can deliver significant speedups for in-order processor designs. Ronald D. Barnes, Erik M. Nystrom, John W. Sias, Sanjay J. Patel, Nacho Navarro, Wen-Mei W. Hwu |
MICRO | 4 |
| 2003 | Characterization of essential dynamic instructionsabstractNo abstract available. Steven S. Lumetta, Sanjay J. Patel |
SIGMETRICS | 2 |
| 2002 | Instruction fetch deferral using static slackabstractIn this paper we present an approach to boosting performance and tolerating latency by deferring non-critical instructions into a deferred queue for later processing. As such, instruction deferral allows more critical instructions to be fetched, dispatched, and possibly executed, earlier. We present methods for identifying deferrable instructions using previously investigated notions of instruction slack. In particular we use static slack to determine if an instruction is deferrable. The static slack of an instruction corresponds to the number of cycles an instruction can be delayed without impacting overall execution time when considering all dynamic paths from that instruction. A significant fraction of the dynamic instruction stream has enough static slack to be deferred by 10 or more cycles on an aggressive execution model. Furthermore, the small amount of register-based communication from deferred instructions to non-deferred instructions makes a deferral-based approach to fetch and execution very attractive. We use a trace cache based microarchitecture to overcome some significant implementation challenges associated with instruction deferral. Overall, instruction deferral boosts the performance of a 4-wide processor by approximately 11% and an 8-wide processor by 6% on eight of the SPEC2000 integer benchmarks. Gregory A. Muthler, David Crowe, Sanjay J. Patel, Steven S. Lumetta |
MICRO | 3 |
| 2001 | Performance characterization of a hardware mechanism for dynamic optimizationabstractWe evaluate the rePLay microarchitecture as a means for reducing application execution time by facilitating dynamic optimization. The framework contains a programmable optimization engine coupled with a hardware-based recovery mechanism. The optimization engine enables the dynamic optimizer to run concurrently with program execution. The recovery mechanism enables the optimizer to make speculative optimizations without requiring recovery code. We demonstrate that a rePLay configuration performing a small suite of simple optimizations on Alpha code attains an average of 13% reduction in execution cycles on the SPEC2000 integer benchmarks over a rePLay configuration not performing optimizations, and a 21% reduction over an aggressive standard superscalar microarchitecture. Brian Fahs, Satarupa Bose, Matthew M. Crum, Brian Slechta, Francesco Spadini, Tony Tung, Sanjay J. Patel, Steven S. Lumetta |
MICRO | 7 |
| 2001 | rePLay: A Hardware Framework for Dynamic OptimizationabstractIn this paper, we propose a new processor framework that supports dynamic optimization. The rePLay Framework embeds an optimization engine atop a high-performance execution engine. The heart of the rePLay Framework is the concept of a frame. Frames are large, single-entry, single-exit optimization regions spanning many basic blocks in the program's dynamic instruction stream, yet containing only a single flow of control. This atomic property of frames increases the flexibility in applying optimizations. To support frames, rePLay includes a hardware-based recovery mechanism that rolls back the architectural state to the beginning of a frame if, for example, an early exit condition is detected. This mechanism permits the optimizer to make speculative, aggressive optimizations upon frames. In this paper, we investigate some of the underlying phenomenon that support rePLay. Primarily, we evaluate rePLay's region formation strategy. A rePLay configuration with a 256-entry frame cache, using 74 KB frame constructor and frame sequencer, achieves an average frame size of 88 Alpha AXP instructions with 68 percent coverage of the dynamic istream, an average frame completion rate of 92.81 percent, and a frame predictor accuracy of 81.26 percent. These results soundly demonstrate that the frames upon which the optimizations are performed are large and stable. Using the most frequently initiated frames from rePLay executions as samples, we also highlight possible strategies for the rePLay optimization engine. Coupled with the high coverage of frames achieved through the dynamic frame construction, the success of these optimizations demonstrates the significance of the rePLay Framework. We believe that the concept of frames, along with the mechanisms and strategies outlined in this paper, will play an important role in future processor architecture. Sanjay J. Patel, Steven S. Lumetta |
IEEE Trans. Computers | 1 |
| 2000 | Increasing the size of atomic instruction blocks using control flow assertionsabstractFor a variety of reasons, branch-less regions of instructions are desirable for high-performance execution. In this paper we propose a means for increasing the dynamic length of branch-less regions of instructions for the purposes of dynamic program optimization. We call these atomic regions frames and we construct them by replacing original branch instructions with assertions. Assertion instructions check if the original branching conditions still hold. If they hold, no action is taken. If they do not, then the entire region is undone. In this manner an assertion has no explicit control flow. We demonstrate that using branch correlation to decide when a branch should be converted into an assertion results in atomic regions that average over 100 instructions in length, with a probability of completion of 97%, and that constitute over 80% of the dynamic instruction stream. We demonstrate both static and dynamic means for constructing frames. When frames are built dynamically using finite sized hardware, they average 80 instructions in length and have good caching properties. Sanjay J. Patel, Tony Tung, Satarupa Bose, Matthew M. Crum |
MICRO | 1 |
| 1999 | Evaluation of Design Options for the Trace Cache Fetch MechanismabstractIn this paper, we examine some critical design features of a trace cache fetch engine for a 16-wide issue processor and evaluate their effects on performance. We evaluate path associativity, partial matching, and inactive issue, all of which are straightforward extensions to the trace cache. We examine features such as the fill unit and branch predictor design. In our final analysis, we show that the trace cache mechanism attains a 28 percent performance improvement over an aggressive single block fetch mechanism and a 15 percent improvement over a sequential multiblock mechanism. Sanjay J. Patel, Daniel H. Friendly, Yale N. Patt |
IEEE Trans. Computers | 1 |
| 1998 | An Analysis of Correlation and Predictability: What Makes Two-Level Branch Predictors WorkabstractPipeline flushes due to branch mispredictions is one of the most serious problems facing the designer of a deeply pipelined, superscalar processor. Many branch predictors have been proposed to help alleviate this problem, including two-level adaptive branch predictors and hybrid branch predictors. Numerous studies have shown which predictors and configurations best predict the branches in a given set of benchmarks. Some studies have also investigated effects, such as pattern history table interference, that can be detrimental to the performance of these predictors. However, little research has been done on which characteristics of branch behavior make predictors perform well. In this paper we investigate and quantify reasons why branches are predictable. We show that some of this predictability is not captured by the two-level adaptive branch predictors. An understanding of the predictability of branches may lead to insights ultimately resulting in better or less complex predictors. We also investigate and quantify what function of the branches in each benchmark is predictable using each of the methods described in this paper. Marius Evers, Sanjay J. Patel, Robert Chappell, Yale N. Patt |
ISCA | 2 |
| 1998 | Improving Trace Cache Effectiveness with Branch Promotion and Trace PackingabstractThe increasing widths of superscalar processors are placing greater demands upon the fetch mechanism. The trace cache meets these demands by placing logically contiguous instructions in physically contiguous storage. As a result, the trace cache delivers instructions at a high rate by supplying multiple fetch blocks each cycle. In this paper we examine two techniques to improve the number of instructions delivered each cycle by the trace cache. The first technique, branch promotion, dynamically converts strongly biased branches into branches with static predictions. Because these promoted branches require no dynamic prediction, the branch predictor suffers less from the negative effects of interference. Branch promotion unlocks the potential of the second technique: trace packing. With trace packing, trace segments are packed with as many instructions as will fit, without regard to naturally occurring fetch block boundaries. With both techniques, the effective fetch rate of the trace cache jumps up 17% over a trace cache which implements neither on a machine where the execution engine has a very aggressive memory disambiguator; the performance of a machine using branch promotion and trace packing is on average 11% higher than a machine using neither technique. Sanjay J. Patel, Marius Evers, Yale N. Patt |
ISCA | 1 |
| 1998 | Putting the Fill Unit to Work: Dynamic Optimizations for Trace Cache MicroprocessorsabstractThe fill unit is the structure which collects blocks of instructions and combines them into multi-block segments for storage in a trace cache. In this paper we expand the role of the fill unit to include four dynamic optimizations: (1) Register move instructions are explicitly marked, enabling them to be executed within the decode logic, (2) Immediate values of dependent instructions are combined, if possible, which removes a step in the dependency chain. (3) Dependent pairs of shift and add instructions are combined into scaled add instructions. (4) Instructions are arranged within the trace segment to minimize the impact of the latency through the operand bypass network. Together these dynamic trace optimizations improve performance on the SPECint95 benchmarks by more than 17% and over all the benchmarks studied by slightly more than 18%. Daniel H. Friendly, Sanjay J. Patel, Yale N. Patt |
MICRO | 2 |
| 1997 | Alternative Fetch and Issue Policies for the Trace Cache Fetch MechanismabstractThe increasing widths of superscalar processors are placing greater demands upon the fetch mechanism. The trace cache meets these demands by placing logically contiguous instructions in physically contiguous storage. It is capable of supplying multiple fetch blocks each cycle. We examine two fetch and issue techniques, partial matching and inactive issue, that improve the overall performance of the trace cache by improving the effective fetch rate. We show that for the SPECint95 benchmarks partial matching increases the overall performance by 12% and adding inactive issue increases performance by 15%. Furthermore we apply these two techniques to issue blocks from trace segments which contain multiple execution paths. We conclude with a performance comparison between a trace cache implementing partial matching and inactive issue and an aggressive single block fetch mechanism. The trace cache increases performance by an average of 25% over the instruction cache. Daniel H. Friendly, Sanjay J. Patel, Yale N. Patt |
MICRO | 2 |
| 1986 | Effectiveness of heuristics measures for automatic test pattern generationabstractThe PODEM (path-oriented decision making) algorithm generates tests for single stuck-at faults in combinational circuits described at the gate level. Controllability and observability (C/O) values were used as heuristic measures in PODEM. The PODEM algorithm and five different methods of determining controllability and observability were implemented. An experiment was conducted to study the effectiveness of different controllability and observability measures for PODEM. Based on the experimental data a new strategy for test generation is discussed. Sanjay J. Patel, Janak H. Patel |
DAC | 1 |