VLDB 2026 Research / reviewers in the wild / expert
Kumar N. Ganapathy
dblp:27/3179
· DBLP profile ↗
4ranked-venue papers
4as first author
0since 2021 · last 1997
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 54% Memory systems · 25% Electronic design automation · 22% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
array processor |
0.0 | 2 | 1997 | Designing a Scalable Processor Array for Recurrent Computations · IEEE Trans. Parallel Distributed Syst. 1997 Optimal Synthesis of Algorithm-Specific Lower-Dimensional Processor Arrays · IEEE Trans. Parallel Distributed Syst. 1996 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.0 | 1 | 1997 | Designing a Scalable Processor Array for Recurrent Computations · IEEE Trans. Parallel Distributed Syst. 1997 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 1996 | Optimal Synthesis of Algorithm-Specific Lower-Dimensional Processor Arrays · IEEE Trans. Parallel Distributed Syst. 1996 |
Methods — techniques the papers use, named apart from their topics
mapping algorithm · 0.0data dependence graph partitioning · 0.0general parameter method · 0.0dependence-based methods · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1997 | Designing a Scalable Processor Array for Recurrent ComputationsabstractIn this paper, we study the design of a coprocessor (CoP) to execute efficiently recursive algorithms with uniform dependencies. Our design is based on two objectives: 1) fixed bandwidth to main memory (MM) and 2) scalability to higher performance without increasing MM bandwidth. Our CoP has an access unit (AU) organized as multiple queues, a processor array (PA) with regularly connected processing elements (PEs), and input/output networks for data routing. Our design is unique because it addresses input/output bottleneck and scalability, two of the most important issues in integrating processor arrays in current systems. To allow processor arrays to be widely usable, they must be scalable to high performance with little or no impact on the supporting memory system. The use of multiple queues in AU also eliminates the use of explicit data addresses, thereby simplifying the design of the control program. We present a mapping algorithm that partitions a data dependence graph (DG) of an application into regular blocks, sequences the blocks through AU, and schedules the execution of the blocks, one at a time, on PA. We show that our mapping procedure minimizes the amount of communication between blocks in the partitioned DG, and sequences the blocks through AU to reduce the communication between AU and MM. Using the matrix-product and transitive-closure applications, we study design trade-offs involving 1) division of a fixed chip area between PA and AU, and 2) improvements in speedup with respect to increases in chip area. Our results show, for a fixed chip area, 1) that there is little degradation in throughput in using a linear PA as compared to a PA organized as a square mesh, and 2) that the design is not sensitive to the division of chip area between PA and AU. We further show that, for a fixed throughput, there is an inverse square root relationship between speedup and total chip area. Our study demonstrates the feasibility of a low-cost, memory bandwidth-limited, and scalable coprocessor system for evaluating recurrent algorithms with uniform dependencies. Kumar N. Ganapathy, Benjamin W. Wah, Chien-Wei Li |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1996 | Optimal Synthesis of Algorithm-Specific Lower-Dimensional Processor ArraysabstractProcessor arrays are frequently used to deliver high performance in many applications with computationally intensive operations. This paper presents the general parameter method (GPM), a systematic parameter-based approach for synthesizing such algorithm-specific architectures. GPM can synthesize processor arrays of any lower dimension from a uniform-recurrence description of the algorithm. The design objective is a general nonlinear and nonmonotonic user-specified function, and depends on attributes such as computation time of the recurrence on the processor array, completion time, load time, and drain time. In addition, bounds on some or all of these attributes can be specified. GPM performs an efficient search of polynomial complexity to find the optimal design satisfying the user-specified design constraints. As an illustration, we show how GPM can be used to find optimal linear processor arrays for computing transitive closures. We consider design objectives that minimize computation time, or processor count, or completion time (including load and drain times), and user-specified constraints on number of processing elements and/or computation/completion times. We show that GPM can be used to obtain optimal designs that trade between number of processing elements and completion time, thereby allowing the designer to choose a design that best meets the specified design objectives. We also show the equivalence between the model assumed in GPM and that in the popular dependence-based methods. Consequently, GPM can be used to find optimal designs for both models. Kumar N. Ganapathy, Benjamin W. Wah |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1992 | Optimal design of lower dimensional processor arrays for uniform recurrencesabstractThe authors present a parameter-based approach for synthesizing systolic architectures from uniform recurrence equations. The scheme presented is a generalization of the parameter method proposed by G.J. Li and B.W. Wah (1985). The approach synthesizes optimal arrays of any lower dimension from a general uniform recurrence description of the problem. In other previous attempts for mapping uniform recurrences into lower-dimensional arrays, optimality of the resulting designs is not guaranteed. As an illustration of the technique, optimal linear arrays for matrix multiplication are given. A detailed design for solving path-finding problems is also presented.> Kumar N. Ganapathy, Benjamin W. Wah |
ASAP | 1 |
| 1992 | Synthesizing Otimal Lower Dimensional Processor Arrays
Kumar N. Ganapathy, Benjamin W. Wah |
ICPP (3) | 1 |