VLDB 2026 Research / reviewers in the wild / expert
G. N. Srinivasa Prasanna
dblp:36/1603 · also Gorur Narayana Srinivasa Prasanna
· DBLP profile ↗
8ranked-venue papers
7as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 76% Parallel and multicore computing · 12% Embedded and real-time systems · 9% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › computer arithmetic
decimal floating-point arithmetic |
0.1 | 1 | 2012 | On Basic Financial Decimal Operations on Binary Machines · IEEE Trans. Computers 2012 |
Processor architecture and microarchitecture
instruction set architecture |
0.1 | 1 | 2012 | On Basic Financial Decimal Operations on Binary Machines · IEEE Trans. Computers 2012 |
Compilers and program optimization
parallelizing compiler |
0.0 | 2 | 1997 | Compilation of Parallel Multimedia Computations - Extending Retiming Theory and Amdahl's Law · PPoPP 1997 Hierarchical Compilation of Macro Dataflow Graphs for Multiprocessors with Local Memory · IEEE Trans. Parallel Distributed Syst. 1994 |
Parallel and multicore computing › task scheduling
DAG scheduling |
0.0 | 2 | 1996 | Generalized Multiprocessor Scheduling and Applications to Matrix Computations · IEEE Trans. Parallel Distributed Syst. 1996 Generalized multiprocessor scheduling for directed acyclic graphs · SC 1994 |
Embedded and real-time systems › real-time scheduling
multiprocessor scheduling |
0.0 | 2 | 1996 | Generalized Multiprocessor Scheduling and Applications to Matrix Computations · IEEE Trans. Parallel Distributed Syst. 1996 Hierarchical Compilation of Macro Dataflow Graphs for Multiprocessors with Local Memory · IEEE Trans. Parallel Distributed Syst. 1994 |
Parallel and multicore computing
processor allocation |
0.0 | 1 | 1994 | Hierarchical Compilation of Macro Dataflow Graphs for Multiprocessors with Local Memory · IEEE Trans. Parallel Distributed Syst. 1994 |
Electronic design automation › high-level synthesis
scheduling |
0.0 | 1 | 1994 | Generalized multiprocessor scheduling for directed acyclic graphs · SC 1994 |
Mathematical optimization › scheduling
multiprocessor scheduling |
0.0 | 1 | 1994 | Generalized multiprocessor scheduling for directed acyclic graphs · SC 1994 |
Mathematical optimization
scheduling |
0.0 | 1 | 1994 | Generalized multiprocessor scheduling for directed acyclic graphs · SC 1994 |
Parallel and multicore computing › multiprocessor system › shared-memory multiprocessor
distributed shared-memory multiprocessor |
0.0 | 2 | 1996 | Generalized Multiprocessor Scheduling and Applications to Matrix Computations · IEEE Trans. Parallel Distributed Syst. 1996 Hierarchical Compilation of Macro Dataflow Graphs for Multiprocessors with Local Memory · IEEE Trans. Parallel Distributed Syst. 1994 |
Embedded and real-time systems
multimedia applications |
0.0 | 1 | 1997 | Compilation of Parallel Multimedia Computations - Extending Retiming Theory and Amdahl's Law · PPoPP 1997 |
Methods — techniques the papers use, named apart from their topics
error sequence analysis · 0.3decimal arithmetic emulation · 0.3optimal control theory · 0.1conjugate gradient minimization · 0.1retiming theory · 0.0amdahl's law · 0.0simulation · 0.0partitioning · 0.0hierarchical compilation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | On Basic Financial Decimal Operations on Binary MachinesabstractFinancial transactions are specified in decimal arithmetic. Until the introduction of IEEE 754-2008, specialized software/hardware routines were used to perform these transactions but it incurred a penalty on performance. In this paper, we show that if binary arithmetic is used to emulate decimal operations, then arbitrary error sequences can be generated by carefully chosen sequences of transactions which can lead to monotonically increasing/decreasing capitalization errors. In addition, we describe methods for correctly performing basic decimal operations, such as addition, subtraction, multiplication, and division, on binary machines, which are not conformant with IEEE 754-2008 decimal floating point standard (ISO/IEC/IEEE 60559:2011), at high speed. Abhilasha Aswal, Ganesh Perumal, G. N. Srinivasa Prasanna |
IEEE Trans. Computers | 3 |
| 1997 | Compilation of Parallel Multimedia Computations - Extending Retiming Theory and Amdahl's LawabstractMultimedia applications (also called multimedia systems) operate on datastreams, which are periodic sequences of data elements, called datasets. A large class of multimedia applications is described by the macro-dataflow graph model, with nodes representing parallelizable tasks, and arcs representing communication. This paper examines how such multimedia applications can be compiled to run efficiently on parallel machines, by optimizing both throughput (T) and latency (L), using two techniques, based on task speedup functions. The first step chooses an appropriate pipeline structure for the system (task clustering). The second step exploits the dataset parallelism intrinsic in the periodic datastream, and runs multiple datasets in parallel (task/cluster multiplicity) for each clustering. The key find-of this research areA The best task clustering depends on system throughput. In general skewed parallelism profiles are desirable i.e. tasks with good speedup and tasks with poor speedup are in separate clusters. Indeed the maximal throughput and minimal latency can be simultaneously attained in the limiting case of a maximally skewed distribution. This result can be viewed as a generalization of Amdahl's law for real-time applications.B Optimal dataset multiplicity for a specific clustering can be determined by extending retiming theory [1] to include parallel resource allocation. In this process, counter-intuitive relaxation regions often appear, wherein by increasing dataset multiplicity, throughput is increased and latency simultaneously reduced (a free lunch).The techniques have been used for compiling real-time image-processing problems on an NCUBE-2 multiprocessor, and show substantial performance gains. G. N. Srinivasa Prasanna |
PPoPP | 1 |
| 1996 | The Optimal Control Approach to Generalized Multiprocessor Scheduling
G. N. Srinivasa Prasanna, Bruce R. Musicus |
Algorithmica | 1 |
| 1996 | Generalized Multiprocessor Scheduling and Applications to Matrix ComputationsabstractThe paper considerably extends the multiprocessor scheduling techniques of G.N.S. Prasanna and B.R. Musicus (1995; 1991) and applies it to matrix arithmetic compilation. Using optimal control theory in the special case where the speedup function of each task is p/sup /spl alpha// (where p is the amount of processing power applied to the task), closed form solution for task graphs formed from parallel and series connections was derived by G.N.S. Prasanna and B.R. Musicus (1995; 1991). The paper extends these results for arbitrary DAGS. The optimality conditions impose nonlinear constraints on the flow of processing power from predecessors to successors, and on the finishing times of siblings. The paper presents a fast algorithm for determining and solving these nonlinear equations. The algorithm utilizes the structure of the finishing time equations to efficiently run a conjugate gradient minimization, leading to the optimal solution. The algorithm has been tested on a variety of DAGs commonly encountered in matrix arithmetic. The results show that if the p/sup /spl alpha// speedup assumption holds, the schedules produced are superior to heuristic approaches. The algorithm has been applied to compiling matrix arithmetic (K.P. Belkhale and P. Banerjee, 1993), for the MIT Alewife machine, a distributed shared memory multiprocessor. While matrix arithmetic tasks do not exactly satisfy the p/sup /spl alpha// speedup assumptions, the algorithm can be applied as a good heuristic. The results show that the schedules produced by our algorithm are faster than alternative heuristic techniques. G. N. Srinivasa Prasanna, Bruce R. Musicus |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1994 | Generalized multiprocessor scheduling for directed acyclic graphsabstractIn the 3rd Annual ACM Symposium on Parallel Algorithms and Architectures, pp. 216-228 (JuIy 1991), we presented several new results in the theory of homogeneous multiprocessor scheduling. A directed acyclic graph (DAG) of tasks was to be scheduled. Tasks were assumed to be parallelizable-as more processors are applied to a task, the time taken to compute it decreases, yielding some speedup. Because of communication, synchronization and task scheduling overheads, this speedup increases less than linearly with the number of processors applied. The optimal scheduling problem is to determine the number of processors assigned to each task, and to the task sequencing, to minimise the finishing time. Using optimal control theory, in the special case where the speedup function of each task is p/sup /spl alpha// (where p is the amount of processing power applied to the task), a closed form solution for task graphs formed from parallel and series connections was derived. This paper considerably extends these techniques for arbitrary DAGs and applies them to matrix arithmetic compilation. The optimality conditions impose nonlinear constraints on the flow of processing power from predecessors to successors, and on the finishing times of siblings. This paper presents a fast algorithm for determining and solving these nonlinear equations. The algorithm utilizes the structure of the finishing time equations to efficiently run a conjugate gradient minimization leading to the optimal solution. The algorithm has been tested on a variety of DAGs. The results presented show that it is superior to alternative heuristic approaches.> G. N. Srinivasa Prasanna, Bruce R. Musicus |
SC | 1 |
| 1994 | Hierarchical Compilation of Macro Dataflow Graphs for Multiprocessors with Local MemoryabstractThis paper presents a hierarchical approach for compiling macro dataflow graphs for multiprocessors with local memory. Macro dataflow graphs comprise several nodes (or macro operations) that must be executed subject to prespecified precedence constraints. Programs consisting of multiple nested loops, where the precedence constraints between the loops are known, can be viewed as macro dataflow graphs. The hierarchical compilation approach comprises a processor allocation phase followed by a partitioning phase. In the processor allocation phase, using estimated speedup functions for the macro nodes, computationally efficient techniques establish the sequencing and parallelism of macro operations for close-to-optimal run-times. The second phase partitions the computations in each macro node to maximize communication locality for the level of parallelism determined by the processor allocation phase. The same approach can also be used for programs consisting of multiple loop nests, when each of the nested loops can be characterized by a speedup function. These ideas have been implemented in a prototype structure-driven compiler, SDC, for expressions of matrix operations. The paper presents the performance of the compiler for several matrix expressions on a simulator of the Alewife multiprocessor.> G. N. Srinivasa Prasanna, Anant Agarwal, Bruce R. Musicus |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1992 | Compile-time Techniques for Processor Allocation in Macro Dataflow Graphs for Multiprocessors
G. N. Srinivasa Prasanna, Anant Agarwal |
ICPP (2) | 1 |
| 1991 | Generalised Multiprocessor Scheduling Using Optimal ControlabstractArticle Generalised multiprocessor scheduling using optimal control Share on Authors: G. N. Srinivasa Prasanna Laboratory for Computer Science and Research Laboratory for Electronics, Massachusetts Institute of Technology, Cambridge, MA Laboratory for Computer Science and Research Laboratory for Electronics, Massachusetts Institute of Technology, Cambridge, MAView Profile , Bruce R. Musicus Laboratory for Computer Science and Research Laboratory for Electronics, Massachusetts Institute of Technology, Cambridge, MA Laboratory for Computer Science and Research Laboratory for Electronics, Massachusetts Institute of Technology, Cambridge, MAView Profile Authors Info & Claims SPAA '91: Proceedings of the third annual ACM symposium on Parallel algorithms and architecturesJune 1991 Pages 216–228https://doi.org/10.1145/113379.113399Published:01 June 1991 18citation374DownloadsMetricsTotal Citations18Total Downloads374Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access G. N. Srinivasa Prasanna, Bruce R. Musicus |
SPAA | 1 |