EDBT 2026 Demo / reviewers in the wild / expert
Rishiyur S. Nikhil
dblp:91/346
· DBLP profile ↗
20ranked-venue papers
6as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 9 · 4 first-authorTheory of computation · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Electronic design automation · 38% Parallel and multicore computing · 37% Processor architecture and microarchitecture · 12% | |
| Software engineering, system software, and programming languages
5 papers |
Runtime systems and virtual machines · 38% Operating systems · 24% Programming languages and type systems · 24% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
system-level design |
0.1 | 1 | 2007 | TLM: Crossing Over From Buzz To Adoption · DAC 2007 |
Electronic design automation › system-level design › high-level system modeling
transaction-level modeling |
0.1 | 1 | 2007 | TLM: Crossing Over From Buzz To Adoption · DAC 2007 |
Parallel and multicore computing
parallel programming models |
0.0 | 4 | 1999 | Space-Time Memory: A Parallel Programming Abstraction for Interactive Multimedia Applications · PPoPP 1999 Exploiting Parallelism in the Implementation of Agna, a Persistent Programming System · ICDE 1991 I-Structures: Data Structures for Parallel Computing · ACM Trans. Program. Lang. Syst. 1989 |
Parallel and multicore computing › parallel programming models and runtimes
parallel programming frameworks |
0.0 | 1 | 2000 | Garbage collection of timestamped data in Stampede · PODC 2000 |
Parallel and multicore computing
task scheduling |
0.0 | 1 | 1999 | Scheduling Constrained Dynamic Applications on Clusters · SC 1999 |
Processor architecture and microarchitecture
dataflow architecture |
0.0 | 3 | 1992 | *T: A Multithreaded Massively Parallel Architecture · ISCA 1992 Executing a Program on the MIT Tagged-Token Dataflow Architecture · IEEE Trans. Computers 1990 Can Dataflow Subsume von Neumann Computing? · ISCA 1989 |
Electronic design automation › hardware verification and test
hardware verification |
0.0 | 1 | 2007 | TLM: Crossing Over From Buzz To Adoption · DAC 2007 |
Integrated circuit design
system-on-chip |
0.0 | 1 | 2007 | TLM: Crossing Over From Buzz To Adoption · DAC 2007 |
High-performance computing
cluster computing |
0.0 | 2 | 2003 | Stampede: A Cluster Programming Middleware for Interactive Stream-Oriented Applications · IEEE Trans. Parallel Distributed Syst. 2003 Garbage collection of timestamped data in Stampede · PODC 2000 |
Processor architecture and microarchitecture
multithreading |
0.0 | 2 | 1992 | *T: A Multithreaded Massively Parallel Architecture · ISCA 1992 Can Dataflow Subsume von Neumann Computing? · ISCA 1989 |
Processor architecture and microarchitecture › dataflow architecture › dataflow machine
dynamic dataflow |
0.0 | 1 | 1992 | *T: A Multithreaded Massively Parallel Architecture · ISCA 1992 |
Parallel and multicore computing › parallel architecture
massively parallel architecture |
0.0 | 1 | 1992 | *T: A Multithreaded Massively Parallel Architecture · ISCA 1992 |
Storage systems › flash and SSD › flash memory management
garbage collection |
0.0 | 1 | 2000 | Garbage collection of timestamped data in Stampede · PODC 2000 |
Operating systems › persistence
persistent object systems |
0.0 | 1 | 1991 | Exploiting Parallelism in the Implementation of Agna, a Persistent Programming System · ICDE 1991 |
Parallel and multicore computing › parallel computation models
data-driven computation |
0.0 | 1 | 1991 | Exploiting Parallelism in the Implementation of Agna, a Persistent Programming System · ICDE 1991 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 1999 | Scheduling Constrained Dynamic Applications on Clusters · SC 1999 |
Parallel and multicore computing
parallel programming runtimes |
0.0 | 1 | 1999 | Space-Time Memory: A Parallel Programming Abstraction for Interactive Multimedia Applications · PPoPP 1999 |
Programming languages and type systems › language semantics › formal semantics
operational semantics |
0.0 | 1 | 1989 | I-Structures: Data Structures for Parallel Computing · ACM Trans. Program. Lang. Syst. 1989 |
Parallel and multicore computing
concurrent data structures |
0.0 | 1 | 1989 | I-Structures: Data Structures for Parallel Computing · ACM Trans. Program. Lang. Syst. 1989 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1989 | Can Dataflow Subsume von Neumann Computing? · ISCA 1989 |
Compilers and program optimization › parallel program optimization
compiler optimization for parallel architectures |
0.0 | 1 | 1992 | *T: A Multithreaded Massively Parallel Architecture · ISCA 1992 |
Query processing and optimization
lazy evaluation |
0.0 | 1 | 1982 | An Implementation Technique for Database Query Languages · ACM Trans. Database Syst. 1982 |
Data models and query languages
query language implementation |
0.0 | 1 | 1982 | An Implementation Technique for Database Query Languages · ACM Trans. Database Syst. 1982 |
Compilers and program optimization › parallel language compilation
dataflow compilation |
0.0 | 1 | 1990 | Executing a Program on the MIT Tagged-Token Dataflow Architecture · IEEE Trans. Computers 1990 |
Programming languages and type systems
functional programming |
0.0 | 1 | 1989 | I-Structures: Data Structures for Parallel Computing · ACM Trans. Program. Lang. Syst. 1989 |
Memory systems › memory architecture
memory system support |
0.0 | 1 | 1989 | Can Dataflow Subsume von Neumann Computing? · ISCA 1989 |
Parallel and multicore computing
multiprocessor system |
0.0 | 1 | 1989 | Can Dataflow Subsume von Neumann Computing? · ISCA 1989 |
Database system architecture and tuning
database interface |
0.0 | 1 | 1982 | An Implementation Technique for Database Query Languages · ACM Trans. Database Syst. 1982 |
Methods — techniques the papers use, named apart from their topics
garbage collection · 0.1data flow graph · 0.1systemc · 0.1optimal scheduling framework · 0.0inter-thread synchronization · 0.0buffer management · 0.0fine-grained parallelism · 0.0P-RISC abstract machine · 0.0dynamic dataflow graph execution · 0.0i-structures · 0.0lazy evaluation · 0.0functional programming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | CLARINET: A quire-enabled RISC-V-based framework for posit arithmetic empiricism
Niraj N. Sharma, Riya Jain, Mohana Madhumita Pokkuluri, Sachin B. Patkar, Rainer Leupers, Rishiyur S. Nikhil, Farhad Merchant |
J. Syst. Archit. | 6 |
| 2009 | Using GPCE principles for hardware systems and accelerators: (bridging the gap to HW design)abstractMoore's Law has precipitated a crisis in the creation of hardware systems (ASICs and FPGAs)-how to design such enormously complex concurrent systems quickly, reliably and affordably? At the same time, portable devices, the energy crisis, and high performance computing present a related challenge-how to move complex and high-performance algorithms from software into hardware (for more speed and/or energy efficiency)? Rishiyur S. Nikhil |
GPCE | 1 |
| 2008 | Hands-on Introduction to Bluespec System Verilog (BSV) (Abstract)abstractBSV is a modern, fully synthesizable design language in which all behavior is expressed with Guarded Atomic Actions (rewrite rules). Rules can be systematically composed from fragments across module boundaries using atomic transactional interfaces. BSV has powerful abstraction mechanisms such as expressive and polymorphic types with overloading and strong static type-checking, full orthogonality (all types are first-class), and Turing-complete static elaboration. Thus, BSV is scalable to large, industrial-strength SoCs even while designs remain highly parameterized and succinct. In this tutorial, you will get a solid technical introduction to BSV and learn how it improves many aspects of modern SoC development: modeling, early SW development, architecture exploration, design, verification, and long-term evolution and maintenance. The lectures will be organized around a few serious examples and we will examine and analyze excerpts of their actual source code. The tutorial is also hands-on: Participants who bring their laptops will receive a non-commercial but full-featured short-term installation of the latest release of BSV (native under Linux, and via a VMWare image for other OSs). During the tutorial you will work with lab exercises tied to the lecture content. After the tutorial you will be able to continue your own exploration with plenty of other examples and lab exercises. Arvind 0001, Rishiyur S. Nikhil |
MEMOCODE | 2 |
| 2007 | TLM: Crossing Over From Buzz To AdoptionabstractTransaction-level modeling --- originally used decades ago in the development of telecommunications network architecture --- is now widely used in SoC design. Why? Because the modern SoC is now so complex that systematic modeling and analysis are required to devise the optimal chip architecture. The architectural model is the essential platform that kick-starts two other key tasks --- verification testbench development and software development. In addition, the interoperability imperatives of SoC design and verification, IP reuse, software development, and system evaluation and integration have driven the replacement of proprietary transaction-level modeling methodologies by a TLM standard that leverages the power of SystemC. The standard --- devised by the Open SystemC Initiative (OSCI) in collaboration with the Open Core Protocol International Partnership (OCP-IP) --- covers the multiple levels of abstraction required for all of the foregoing tasks. Francine Bacchini, Daniel Gajski, Laurent Maillet-Contoz, Haruhisa Kashiwagi, Jack Donovan, Tommi Mäkeläinen, Jack Greenbaum, Rishiyur S. Nikhil |
DAC | 8 |
| 2006 | A rule-based model of computation for SystemC: integrating SystemC and Bluespec for co-designabstractBluespec's rule-based model of computation (MoC) for hardware concurrency has gained attention for several reasons. From its basis in term rewriting systems, rules have the property of atomicity, which improves correctness by construction, particularly in large-scale concurrency with finegrained, dynamic resource sharing (typical in complex hardware). Rule-based interface methods extend atomicity across module boundaries, have a natural transactional reading, and precisely and formally characterize resource-sharing constraints. All this can be synthesized to hardware with competitive quality. SystemC expresses concurrency with threading and events, just like RTL, where it is difficult to deal with fine-grain concurrency and resource sharing. Further, there is no systematic methodology for module composition. Thus, while SystemC is suitable for very coarse modeling and for embedded software development, its limitations make it difficult to model correct by construction hardware systems accurately. In this paper, we show how to integrate Bluespec's rule-based MoC into SystemC. We augment SystemC modules with rules and rule-based interface methods, and augment the SystemC simulation kernel with a rule execution kernel. The integration is augmentative in that a model can contain both rule-based modules (where hardware accuracy is desired) as well as core SystemC or TLM modules (for embedded software, instruction-set simulators, existing SystemC IP, or pure behavioral models), thus providing the advantages of each MoC where appropriate Hiren D. Patel, Sandeep K. Shukla, Elliot Mednick, Rishiyur S. Nikhil |
MEMOCODE | 4 |
| 2005 | Synthesis of synchronous assertions with guarded atomic actionsabstractThe SystemVerilog standard introduces SystemVerilog Assertions (SVA), a synchronous assertion package based on the temporal-logic semantics of PSL. Traditionally assertions are checked in software simulation. We introduce a method for synthesizing SVA directly into hardware modules in Bluespec SystemVerilog. This opens up new possibilities for FPGA-accelerated testbenches, hardware/software co-emulation, dynamic verification and fault-tolerance. We describe adding synthesizable assertions to a cache controller, and investigate their hardware cost. Michael Pellauer, Mieszko Lis, Don Baltus, Rishiyur S. Nikhil |
MEMOCODE | 4 |
| 2004 | High-level synthesis: an essential ingredient for designing complex ASICsabstractIt is common wisdom that synthesizing hardware from higher-level descriptions than Verilog incurs a performance penalty. The case study here shows that this need not be the case. If the higher-level language has suitable semantics, it is possible to synthesize hardware that is competitive with hand-written Verilog RTL. Differences in the hardware quality are dominated by architecture differences and, therefore, it is more important to explore multiple hardware architectures. This exploration is not practical without quality synthesis from higher-level languages. Arvind 0001, Rishiyur S. Nikhil, Daniel L. Rosenband, Nirav Dave |
ICCAD | 2 |
| 2004 | Bluespec System Verilog: efficient, correct RTL from high level specificationsabstractBluespec System Verilog is an EDL toolset for ASIC and FPGA design offering significantly higher productivity via a radically different approach to high-level synthesis. Many other attempts at high-level synthesis have tried to move the design language towards a more software-like specification of the behavior of the intended hardware. By means of code samples, demonstrations and measured results, we illustrate how Bluespec System Verilog, in an environment familiar to hardware designers, can significantly improve productivity without compromising generated hardware quality. Rishiyur S. Nikhil |
MEMOCODE | 1 |
| 2003 | Stampede: A Cluster Programming Middleware for Interactive Stream-Oriented ApplicationsabstractEmerging application domains such as interactive vision, animation, and multimedia collaboration display dynamic scalable parallelism and high-computational requirements, making them good candidates for executing on parallel architectures such as SMPs and clusters of SMPs. Stampede is a programming system that has many of the needed functionalities such as high-level data sharing, dynamic cluster-wide threads and their synchronization, support for task and data parallelism, handling of time-sequenced data items, and automatic buffer management. We present an overview of Stampede, the primary data abstractions, the algorithmic basis of garbage collection, and the issues in implementing these abstractions on a cluster of SMPs. We also present a set of micromeasurements along with two multimedia applications implemented on top of Stampede, through which we demonstrate the low overhead of this runtime and that it is suitable for the streaming multimedia applications. Umakishore Ramachandran, Rishiyur S. Nikhil, James M. Rehg, Yavor Angelov, Arnab Paul, Sameer Adhikari, Kenneth M. Mackenzie, Nissim Harel, Kathleen Knobe |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2000 | Garbage collection of timestamped data in StampedeabstractStampede is a parallel programming system to facilitate the programming of interactive multimedia applications on clusters of SMPs. Rishiyur S. Nikhil, Umakishore Ramachandran |
PODC | 1 |
| 1999 | Space-Time Memory: A Parallel Programming Abstraction for Interactive Multimedia ApplicationsabstractRealistic interactive multimedia involving vision, animation, and multimedia collaboration is likely to become an important aspect of future computer applications. The scalable parallelism inherent in such applications coupled with their computational demands make them ideal candidates for SMPs and clusters of SMPs. These applications have novel requirements that offer new kinds of challenges for parallel system design.We have designed a programming system called Stampede that offers many functionalities needed to simplify development of such applications (such as high-level data sharing abstractions, dynamic cluster-wide threads, and multiple address spaces). We have built Stampede and it runs on clusters of SMPs. To date we have implemented two applications on Stampede, one of which is discussed herein.In this paper we describe a part of Stampede called Space-Time Memory (STM). It is a novel data sharing abstraction that enables interactive multimedia applications to manage a collection of time-sequenced data items simply, efficiently, and transparently across a cluster. STM relieves the application programmer from low level synchronization and data communication by providing a high level interface that subsumes buffer management, inter-thread synchronization, and location transparency for data produced and accessed anywhere in the cluster. STM also automatically handles garbage collection of data items that will no longer be accessed by any of the application threads. We discuss ease of use issues for developing applications using STM, and present preliminary performance results to show that STM's overhead is low. Umakishore Ramachandran, Rishiyur S. Nikhil, Nissim Harel, James M. Rehg, Kathleen Knobe |
PPoPP | 2 |
| 1999 | Scheduling Constrained Dynamic Applications on ClustersabstractThere is an emerging class of computationally demanding multimedia applications involving vision, speech and interaction with the real world (e.g., CRL's Smart Kiosk). These applications are highly parallel and require low latencies for good performance. They are well-suited for implementation on clusters of SMP's, but they require efficient scheduling of application tasks. General purpose schedulers produce high latencies because they lack knowledge of the dependencies between tasks. Previous research in optimal scheduling has been limited to static problems. In contrast, our application is highly dynamic as the optimal schedule depends upon the behavior of the kiosk's customers. We observe that the dynamism of our application class is constrained, in that there are a small number of operating regimes which are determined by the state of the application. We present a framework for optimal scheduling of constrained dynamic applications. The results of an experimental compariso... Kathleen Knobe, James M. Rehg, Arun Chauhan 0001, Rishiyur S. Nikhil, Umakishore Ramachandran |
SC | 4 |
| 1996 | A Lambda Calculus with Letrecs and Barriers
Arvind 0001, Jan-Willem Maessen, Rishiyur S. Nikhil, Joseph E. Stoy |
FSTTCS | 3 |
| 1996 | pHluid: The Design of a Parallel Functional Language Implementation on WorkstationsabstractThis paper describes the distributed memory implementation of a shared memory parallel functional language. The language is Id, an implicitly parallel, mostly functional language that is currently evolving into a dialect of Haskell. The target is a distributed memory machine, because we expect these to be the most widely available parallel platforms in the future. The difficult problem is to bridge the gap between the shared memory language model and the distributed memory machine model. The language model assumes that all data is uniformly accessible, whereas the machine has a severe memory hierarchy: a processor's access to remote memory (using explicit communication) is orders of magnitude slower than its access to local memory. Thus, avoiding communication is crucial for good performance. The Id language, and its general dataflow-inspierd compilation to multithreaded code are described elsewhere. In this paper, we focus on our new parallel runtime system and its features for avoiding communication and for tolerating its latency when necessary: multithreading, scheduling and load balancing; the distributed heap model and distributed coherent cacheing, and parallel garbage collection. We have completed the first implementation, and we present some preliminary performance mearsurements. Cormac Flanagan, Rishiyur S. Nikhil |
ICFP | 2 |
| 1992 | *T: A Multithreaded Massively Parallel ArchitectureabstractWhat should the architecture of each node in a general purpose, massively parallel architecture (MPA) be? We frame the question in concrete terms by describing two fundamental problems that must be solved well in any general purpose MPA. From this, we systematically develop the required logical organization of an MPA node, and present some details of *T (pronounced Start, a concrete architecture designed to these requirements. *T is a direct descendant of dynamic dataflow architectures, and unifies them with von Neumann architectures. We discuss a hand-compiled example and some compilation issues. Rishiyur S. Nikhil, Gregory M. Papadopoulos, Arvind 0001 |
ISCA | 1 |
| 1991 | Exploiting Parallelism in the Implementation of Agna, a Persistent Programming SystemabstractA design for AGNA, a persistent object system that utilizes parallelism in a fundamental way to enhance performance, is presented. The underlying thesis is that fine-grained parallelism is essential for achieving scalable performance on parallel multiple instruction/multiple data (MIMD) machines. This, in turn, implies a data-driven model of computation for efficiency. The complete design based on these principles starts with a declarative source language because such languages reveal the most fine-grained parallelism. It is described how transactions are compiled into an abstract, fine-grained parallel machine called P-RISC. The P-RISC virtual heap is implemented in the memory and disk of a parallel machine in such a way that paging is overlapped with useful computation. The current implementation status is described, some preliminary performance results are reported and the approach presented is compared to several recent parallel database system projects.> Rishiyur S. Nikhil, Michael L. Heytens |
ICDE | 1 |
| 1990 | Executing a Program on the MIT Tagged-Token Dataflow ArchitectureabstractThe MIT Tagged-Token Dataflow Project has an unconventional, but integrated approach to general-purpose high-performance parallel computing. Rather than extending conventional sequential languages, Id, a high-level language with fine-grained parallelism and determinacy implicit in its operational semantics, is used. Id programs are compiled to dynamic dataflow graphs, which constitute a parallel machine language. Dataflow graphs are directly executed on the MIT tagged-token dataglow architecture (TTDA), a multiprocessor architecture. An overview of current thinking on dataflow architecture is provided by describing example Id programs, their compilation to dataflow graphs, and their execution on the TTDA. Related work and the status of the project are described.> Arvind 0001, Rishiyur S. Nikhil |
IEEE Trans. Computers | 2 |
| 1989 | Can Dataflow Subsume von Neumann Computing?abstractWe explore the question: “What can a von Neumann processor borrow from dataflow to make it more suitable for a multiprocessor?” Starting with a simple, “RISC-like” instruction set, we show how to change the underlying processor organization to make it multithreaded. Then, we extend it with three instructions that give it a fine-grained, dataflow capability. We call the result P-RISC, for “Parallel RISC.” Finally, we discuss memory support for such multiprocessors. We compare our approach to existing MIMD machines and to other dataflow machines. Rishiyur S. Nikhil |
ISCA | 1 |
| 1989 | I-Structures: Data Structures for Parallel ComputingabstractIt is difficult to achieve elegance, efficiency, and parallelism simultaneously in functional programs that manipulate large data structures. We demonstrate this through careful analysis of program examples using three common functional data-structuring approaches-lists using Cons, arrays using Update (both fine-grained operators), and arrays using make-array (a “bulk” operator). We then present I-structure as an alternative and show elegant, efficient, and parallel solutions for the program examples in Id, a language with I-structures. The parallelism in Id is made precise by means of an operational semantics for Id as a parallel reduction system. I-structures make the language nonfunctional, but do not lose determinacy. Finally, we show that even in the context of purely functional languages, I-structures are invaluable for implementing functional data abstractions. Arvind 0001, Rishiyur S. Nikhil, Keshav Pingali |
ACM Trans. Program. Lang. Syst. | 2 |
| 1982 | An Implementation Technique for Database Query LanguagesabstractStructured query languages, such as those available for relational databases, are becoming increasingly desirable for all database management systems. Such languages are applicative: there is no need for an assignment or update statement. A new technique is described that allows for the implementation of applicative query languages against most commonly used database systems. The technique involves “lazy” evaluation and has a number of advantages over existing methods: it allows queries and functions of arbitrary complexity to be constructed; it reduces the use of secondary storage; it provides a simple control structure through which interfaces to other programs may be constructed; and the implementation, including the database interface, is quite compact. Although the technique is presented for a specific functional programming system and for a CODASYL DBMS, it is general and may be used for other query languages and database systems. Peter Buneman, Robert E. Frankel, Rishiyur S. Nikhil |
ACM Trans. Database Syst. | 3 |