VLDB 2026 Research / reviewers in the wild / expert
Sriram Vajapeyam
dblp:95/5676
· DBLP profile ↗
11ranked-venue papers
7as first author
0since 2021 · last 2000
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 7 first-authorSoftware engineering, systems software and programming languages · 5 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 88% High-performance computing · 6% Performance modeling and evaluation · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.0 | 2 | 1999 | Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999 Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 2 | 1999 | Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999 On the instruction-level characteristics of scalar code in highly-vectorized scientific applications · MICRO 1992 |
Processor architecture and microarchitecture › vector processing
dynamic vectorization |
0.0 | 2 | 1999 | Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999 Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997 |
Processor architecture and microarchitecture › out-of-order execution
register renaming |
0.0 | 1 | 1997 | Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1997 | Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code Sequences · ISCA 1997 |
High-performance computing › supercomputing
supercomputing systems |
0.0 | 1 | 1991 | An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club Benchmarks · ISCA 1991 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1991 | An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club Benchmarks · ISCA 1991 |
Compilers and program optimization
vectorization |
0.0 | 1 | 1999 | Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary Programs · ISCA 1999 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1989 | Tradeoffs in Instruction Format Design for Horizontal Architectures · ASPLOS 1989 |
Processor architecture and microarchitecture
instruction issue logic |
0.0 | 1 | 1987 | Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 1987 | Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987 |
High-performance computing
scientific computing |
0.0 | 1 | 1992 | On the instruction-level characteristics of scalar code in highly-vectorized scientific applications · MICRO 1992 |
Processor architecture and microarchitecture › microprocessor design › processor core design
functional units |
0.0 | 1 | 1987 | Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987 |
Processor architecture and microarchitecture
multiple functional units |
0.0 | 1 | 1987 | Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors · ISCA 1987 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.1empirical study · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2000 | Non-Strict Cache Coherence: Exploiting Data-Race Tolerance in Emerging ApplicationsabstractSoftware distributed shared memory (DSM) platforms on networks of workstations tolerate large network latencies by employing one of several weak memory consistency models. Data-race tolerant applications, such as Genetic Algorithms (GAs), Probabilistic Inference, etc., offer an additional degree of freedom to tolerate network latency: they do not synchronize shared memory references, and behave correctly when supplied outdated shared data. However, these algorithms often have a high communication-to-computation ratio and can flood the network with messages in the presence of large message delays. We study the performance of controlled asynchronous implementations of these algorithms via the use of our previously proposed blocking Global Read memory access primitive. Global Read implements non-strict cache coherence by guaranteeing to return to the reader a shared datum value from within a specified staleness range. Experiments on an IBM SP2 multicomputer with an Ethernet show significant performance improvements for controlled asynchronous implementations. On a lightly loaded Ethernet network, most of the GA benchmarks see 30% to 40% improvement over the best competitor for 2 to 16 processors, while two of the Probabilistic Inference benchmarks see more than 80% improvement for 2 processors. As the network load increases, the benefits of non-strict cache coherence increase significantly. Siddhartha V. Tambat, Sriram Vajapeyam |
ICPP | 2 |
| 1999 | Whither Indian Computer Science R & D?
Sriram Vajapeyam |
HiPC | 1 |
| 1999 | Dynamic Vectorization: A Mechanism for Exploiting Far-Flung ILP in Ordinary ProgramsabstractSeveral ILP limit studies indicate the presence of considerable ILP across dynamically far-apart instructions in program execution. This paper proposes a hardware mechanism, dynamic vectorization (DV), as a tool for quickly building up a large logical instruction window. Dynamic vectorization converts repetitive dynamic instruction sequences into vector form, enabling the processing of instructions from beyond the corresponding program loop to be overlapped with the loop. This enables vector-like execution of programs with relatively complex static control flow that may not be amenable to static, compile time vectorization. Experimental evaluation shows that a large fraction of the dynamic instructions of four of the six SPECInt92 programs can be captured in vector form. Three of these programs exhibit significant potential for ILP improvements from dynamic vectorization, with speedups of more than a factor of 2 in a scenario of realistic branch prediction and perfect memory disambiguation. Under perfect branch prediction conditions, a fourth program also shows well over a factor of 2 speedup from DV. The speedups are due to the overlap of post-loop processing with loop processing. Sriram Vajapeyam, P. J. Joseph, Tulika Mitra |
ISCA | 1 |
| 1997 | Improving Superscalar Instruction Dispatch and Issue by Exploiting Dynamic Code SequencesabstractSuperscalar processors currently have the potential to fetch multiple basic blocks per cycle by employing one of several recently proposed instruction fetch mechanisms. However, this increased fetch bandwidth cannot be exploited unless pipeline stages further downstream correspondingly improve. In particular, register renaming a large number of instructions per cycle is difficult. A large instruction window, needed to receive multiple basic blocks per cycle, will slow down dependence resolution and instruction issue. This paper addresses these and related issues by proposing (i) partitioning of the instruction window into multiple blocks, each holding a dynamic code sequence; (ii) logical partitioning of the register file into a global file and several local files, the latter holding registers local to a dynamic code sequence; (iii) the dynamic recording and reuse of register renaming information for registers local to a dynamic code sequence. Performance studies show these mechanisms improve performance over traditional superscalar processors by factors ranging from 1.5 to a little over 3 for the SPEC Integer programs. Next, it is observed that several of the loops in the benchmarks display vector-like behavior during execution, even if the static loop bodies are likely complex for compile-time vectorization. A dynamic loop vectorization mechanism that builds on top of the above mechanisms is briefly outlined. The mechanism vectorizes up to 60% of the dynamic instructions for some programs, albeit the average number of iterations per loop is quite small. Sriram Vajapeyam, Tulika Mitra |
ISCA | 1 |
| 1996 | Program-level control of network delay for parallel asynchronous iterative applicationsabstractSoftware distributed shared memory (DSM) platforms on networks of workstations tolerate large network latencies by employing one of several weak memory consistency models. Fully asynchronous parallel iterative algorithms offer an additional degree of freedom to tolerate network latency. They behave correctly when supplied outdated shared data. However these algorithms can flood the network with messages in the presence of large delays. We propose a method of controlling asynchronous iterative methods wherein the reader of a shared datum imposes an upper bound on its age via use of a blocking Global Read primitive. This reduces the overall number of iteration is executed by the reader; thus controlling the amount of shared updates generated. Experiments for a fully asynchronous linear equation solver running on a network of 10 IBM RS/6000 workstations show that the proposed Global Read primitive provides significant performance improvement. P. J. Joseph, Sriram Vajapeyam |
HiPC | 2 |
| 1993 | Toward Effective Scalar Hardware for Highly Vectorizable Applications
Sriram Vajapeyam, Wei-Chung Hsu |
J. Parallel Distributed Comput. | 1 |
| 1992 | On the instruction-level characteristics of scalar code in highly-vectorized scientific applications
Sriram Vajapeyam, Wei-Chung Hsu |
MICRO | 1 |
| 1991 | An Empirical Study of the CRAY Y-MP Processor Using the Perfect Club BenchmarksabstractCharacterization of machines, by studying pro~am usage of their architectural and organizational features, IS art essential ~art of the desi~n recess.ln this aper we re ort Y EL an empimcal study of a smg e processor of t e CRAY Y-P, using as benchmarks long-running scientific applications Sriram Vajapeyam, Gurindar S. Sohi, Wei-Chung Hsu |
ISCA | 1 |
| 1990 | Exploitation of operation-level parallelism in a processor of the CRAY X-MPabstractAvailable operation-level parallelism and its exploitation in the CRAY X-MP processor are studied. Considered are the sizes and contributions to execution time of basic blocks, instruction and operation issue rates and issue stalls, and operation execution overlap for entire executions of three large programs, FLO52, TRFD, and QCD1, taken from the Perfect Club benchmark set. The large basic blocks account for a significant portion of the overall execution time. It is also found that with the use of vector instructions, the X-MP is able to issue more than one operation per clock cycle, even though it can issue a maximum of one instruction per cycle.> Sriram Vajapeyam, Gurindar S. Sohi, Wei-Chung Hsu |
ICCD | 1 |
| 1989 | Tradeoffs in Instruction Format Design for Horizontal ArchitecturesabstractWith recent improvements in software techniques and the enhanced level of fine grain parallelism made available by such techniques, there has been an increased interest in horizontal architectures and large instruction words that are capable of issuing more that one operation per instruction. This paper investigates some issues in the design of such instruction formats. We study how the choice of an instruction format is influenced by factors such as the degree of pipelining and the instruction's view of the register file. Our results suggest that very large instruction words capable of issuing one operation to each functional unit resource in a horizontal architecture may be overkill. Restricted instruction formats with limited operation issuing capabilities are capable of providing similar performance (measured by the total number of time steps) with significantly less hardware in many cases. Gurindar S. Sohi, Sriram Vajapeyam |
ASPLOS | 2 |
| 1987 | Instruction Issue Logic for High-Performance, Interruptable Pipelined ProcessorsabstractThe performance of pipelined processors is severely limited by data dependencies. In order to achieve high performance, a mechanism to alleviate the effects of data dependencies must exist. If a pipelined CPU with multiple functional units is to be used in the presence of a virtual memory hierarchy, a mechanism must also exist for determining the state of the machine precisely. In this paper, we combine the issues of dependency-resolution and preciseness of state. We present a design for instruction issue logic that resolves dependencies dynamically and, at the same time, guarantees a precise state of the machine, without a significant hardware overhead. Detailed simulation studies for the proposed mechanism, using the Lawrence Livermore loops as a benchmark, are presented. Gurindar S. Sohi, Sriram Vajapeyam |
ISCA | 2 |