Sriram Vajapeyam

dblp:95/5676 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 7 first-authorSoftware engineering, systems software and programming languages · 5 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Processor architecture and microarchitecture · 88% High-performance computing · 6% Performance modeling and evaluation · 6%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › out-of-order execution
instruction window
0.021999
Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999
Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997
Processor architecture and microarchitecture
instruction-level parallelism
0.021999
Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999
On the instruction-level characteristics of scalar code in highly-vectorized scientific applications · MICRO 1992
Processor architecture and microarchitecture › vector processing
dynamic vectorization
0.021999
Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999
Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.011997
Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997
Processor architecture and microarchitecture
superscalar processor
0.011997
Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997
High-performance computing › supercomputing
supercomputing systems
0.011991
An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club Benchmarks · ISCA 1991
Performance modeling and evaluation
workload characterization
0.011991
An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club Benchmarks · ISCA 1991
Compilers and program optimization
vectorization
0.011999
Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999
Processor architecture and microarchitecture
instruction set architecture
0.011989
Tradeoffs in Instruction Format Design for Horizontal Architectures · ASPLOS 1989
Processor architecture and microarchitecture
instruction issue logic
0.011987
Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987
Processor architecture and microarchitecture
pipelining
0.011987
Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987
High-performance computing
scientific computing
0.011992
On the instruction-level characteristics of scalar code in highly-vectorized scientific applications · MICRO 1992
Processor architecture and microarchitecture › microprocessor design › processor core design
functional units
0.011987
Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987
Processor architecture and microarchitecture
multiple functional units
0.011987
Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987

Methods — techniques the papers use, named apart from their topics

simulation · 0.1empirical study · 0.0
YearPublicationVenuePosition
2000 Non-Strict Cache Coherence: Exploiting Data-Race Tolerance in Emerging Applications
abstract
Software distributed shared memory (DSM) platforms on networks of workstations tolerate large network latencies by employing one of several weak memory consistency models. Data-race tolerant applications, such as Genetic Algorithms (GAs), Probabilistic Inference, etc., offer an additional degree of freedom to tolerate network latency: they do not synchronize shared memory references, and behave correctly when supplied outdated shared data. However, these algorithms often have a high communication-to-computation ratio and can flood the network with messages in the presence of large message delays. We study the performance of controlled asynchronous implementations of these algorithms via the use of our previously proposed blocking Global Read memory access primitive. Global Read implements non-strict cache coherence by guaranteeing to return to the reader a shared datum value from within a specified staleness range. Experiments on an IBM SP2 multicomputer with an Ethernet show significant performance improvements for controlled asynchronous implementations. On a lightly loaded Ethernet network, most of the GA benchmarks see 30% to 40% improvement over the best competitor for 2 to 16 processors, while two of the Probabilistic Inference benchmarks see more than 80% improvement for 2 processors. As the network load increases, the benefits of non-strict cache coherence increase significantly.
Siddhartha V. Tambat, Sriram Vajapeyam
ICPP2
1999 Whither Indian Computer Science R & D?
Sriram Vajapeyam
HiPC1
1999 Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs
abstract
Several ILP limit studies indicate the presence of considerable ILP across dynamically far-apart instructions in program execution. This paper proposes a hardware mechanism, dynamic vectorization (DV), as a tool for quickly building up a large logical instruction window. Dynamic vectorization converts repetitive dynamic instruction sequences into vector form, enabling the processing of instructions from beyond the corresponding program loop to be overlapped with the loop. This enables vector-like execution of programs with relatively complex static control flow that may not be amenable to static, compile time vectorization. Experimental evaluation shows that a large fraction of the dynamic instructions of four of the six SPECInt92 programs can be captured in vector form. Three of these programs exhibit significant potential for ILP improvements from dynamic vectorization, with speedups of more than a factor of 2 in a scenario of realistic branch prediction and perfect memory disambiguation. Under perfect branch prediction conditions, a fourth program also shows well over a factor of 2 speedup from DV. The speedups are due to the overlap of post-loop processing with loop processing.
Sriram Vajapeyam, P. J. Joseph, Tulika Mitra
ISCA1
1997 Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences
abstract
Superscalar processors currently have the potential to fetch multiple basic blocks per cycle by employing one of several recently proposed instruction fetch mechanisms. However, this increased fetch bandwidth cannot be exploited unless pipeline stages further downstream correspondingly improve. In particular, register renaming a large number of instructions per cycle is difficult. A large instruction window, needed to receive multiple basic blocks per cycle, will slow down dependence resolution and instruction issue. This paper addresses these and related issues by proposing (i) partitioning of the instruction window into multiple blocks, each holding a dynamic code sequence; (ii) logical partitioning of the register file into a global file and several local files, the latter holding registers local to a dynamic code sequence; (iii) the dynamic recording and reuse of register renaming information for registers local to a dynamic code sequence. Performance studies show these mechanisms improve performance over traditional superscalar processors by factors ranging from 1.5 to a little over 3 for the SPEC Integer programs. Next, it is observed that several of the loops in the benchmarks display vector-like behavior during execution, even if the static loop bodies are likely complex for compile-time vectorization. A dynamic loop vectorization mechanism that builds on top of the above mechanisms is briefly outlined. The mechanism vectorizes up to 60% of the dynamic instructions for some programs, albeit the average number of iterations per loop is quite small.
Sriram Vajapeyam, Tulika Mitra
ISCA1
1996 Program-level control of network delay for parallel asynchronous iterative applications
abstract
Software distributed shared memory (DSM) platforms on networks of workstations tolerate large network latencies by employing one of several weak memory consistency models. Fully asynchronous parallel iterative algorithms offer an additional degree of freedom to tolerate network latency. They behave correctly when supplied outdated shared data. However these algorithms can flood the network with messages in the presence of large delays. We propose a method of controlling asynchronous iterative methods wherein the reader of a shared datum imposes an upper bound on its age via use of a blocking Global Read primitive. This reduces the overall number of iteration is executed by the reader; thus controlling the amount of shared updates generated. Experiments for a fully asynchronous linear equation solver running on a network of 10 IBM RS/6000 workstations show that the proposed Global Read primitive provides significant performance improvement.
P. J. Joseph, Sriram Vajapeyam
HiPC2
1993 Toward Effective Scalar Hardware for Highly Vectorizable Applications
Sriram Vajapeyam, Wei-Chung Hsu
J. Parallel Distributed Comput.1
1992 On the instruction-level characteristics of scalar code in highly-vectorized scientific applications
Sriram Vajapeyam, Wei-Chung Hsu
MICRO1
1991 An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club Benchmarks
abstract
Characterization of machines, by studying pro~am usage of their architectural and organizational features, IS art essential ~art of the desi~n recess.ln this aper we re ort Y EL an empimcal study of a smg e processor of t e CRAY Y-P, using as benchmarks long-running scientific applications
Sriram Vajapeyam, Gurindar S. Sohi, Wei-Chung Hsu
ISCA1
1990 Exploitation of operation-level parallelism in a processor of the CRAY X-MP
abstract
Available operation-level parallelism and its exploitation in the CRAY X-MP processor are studied. Considered are the sizes and contributions to execution time of basic blocks, instruction and operation issue rates and issue stalls, and operation execution overlap for entire executions of three large programs, FLO52, TRFD, and QCD1, taken from the Perfect Club benchmark set. The large basic blocks account for a significant portion of the overall execution time. It is also found that with the use of vector instructions, the X-MP is able to issue more than one operation per clock cycle, even though it can issue a maximum of one instruction per cycle.>
Sriram Vajapeyam, Gurindar S. Sohi, Wei-Chung Hsu
ICCD1
1989 Tradeoffs in Instruction Format Design for Horizontal Architectures
abstract
With recent improvements in software techniques and the enhanced level of fine grain parallelism made available by such techniques, there has been an increased interest in horizontal architectures and large instruction words that are capable of issuing more that one operation per instruction. This paper investigates some issues in the design of such instruction formats. We study how the choice of an instruction format is influenced by factors such as the degree of pipelining and the instruction's view of the register file. Our results suggest that very large instruction words capable of issuing one operation to each functional unit resource in a horizontal architecture may be overkill. Restricted instruction formats with limited operation issuing capabilities are capable of providing similar performance (measured by the total number of time steps) with significantly less hardware in many cases.
Gurindar S. Sohi, Sriram Vajapeyam
ASPLOS2
1987 Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors
abstract
The performance of pipelined processors is severely limited by data dependencies. In order to achieve high performance, a mechanism to alleviate the effects of data dependencies must exist. If a pipelined CPU with multiple functional units is to be used in the presence of a virtual memory hierarchy, a mechanism must also exist for determining the state of the machine precisely. In this paper, we combine the issues of dependency-resolution and preciseness of state. We present a design for instruction issue logic that resolves dependencies dynamically and, at the same time, guarantees a precise state of the machine, without a significant hardware overhead. Detailed simulation studies for the proposed mechanism, using the Lawrence Livermore loops as a benchmark, are presented.
Gurindar S. Sohi, Sriram Vajapeyam
ISCA2