EDBT 2026 Demo / reviewers in the wild / expert
Sailesh K. Rao
dblp:55/634
· DBLP profile ↗
10ranked-venue papers
5as first author
0since 2021 · last 1997
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-authorTheory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 32% High-performance computing · 19% Hardware accelerators and domain-specific architectures · 18% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
systolic array |
0.0 | 2 | 1990 | On the Analysis of Synchronous Computing Systems · SIAM J. Comput. 1990 Array architectures for iterative algorithms · Proc. IEEE 1987 |
Electronic design automation › system-level design › model of computation
regular iterative algorithm |
0.0 | 2 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 Array architectures for iterative algorithms · Proc. IEEE 1987 |
Parallel and multicore computing
parallel programming models and runtimes |
0.0 | 1 | 1989 | Communication reduction for distributed sparse matrix factorization on a processor mesh · SC 1989 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1989 | Communication reduction for distributed sparse matrix factorization on a processor mesh · SC 1989 |
High-performance computing › sparse linear algebra
sparse matrix factorization |
0.0 | 1 | 1989 | Communication reduction for distributed sparse matrix factorization on a processor mesh · SC 1989 |
Parallel and multicore computing
parallel architecture |
0.0 | 1 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 |
Processor architecture and microarchitecture › parallel computer organization
systolic array design |
0.0 | 1 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 |
Electronic design automation › design methodology
hardware design methodology |
0.0 | 1 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 |
Parallel and multicore computing › task allocation
processor array mapping |
0.0 | 1 | 1988 | Regular iterative algorithms and their implementation on processor arrays · Proc. IEEE 1988 |
Parallel and multicore computing › array processor
mesh-connected processor array |
0.0 | 1 | 1987 | Array architectures for iterative algorithms · Proc. IEEE 1987 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
parallel algorithm mapping |
0.0 | 1 | 1987 | Array architectures for iterative algorithms · Proc. IEEE 1987 |
Methods — techniques the papers use, named apart from their topics
system theory · 0.0graph theory · 0.0matrix permutation · 0.0fragmented distribution · 0.0elimination tree · 0.0systolic array methodology · 0.0dependence mapping · 0.0systolic array design · 0.0algorithm-to-array transformation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1997 | 100BASE-T2: 100 Mbit/s Ethernet over Two Pairs of Category-3 Cablingabstract100BASE-T2 is a new physical layer specification for IEEE 802.3 LANs operating at 100 Mbit/s ("fast Ethernet"). It enables users of the 10BASE-T Ethernet LAN technology to upgrade their networks from 10 Mbit/s to 100 Mbit/s performance while retaining an existing voice-grade cabling infrastructure. 100BASE-T2 transceivers will operate over two pairs in unshielded twisted-pair cables corresponding to EIA/TIA category 3 (UTP-3), which is the minimum requirement for 10BASE-T. In a four-pair UTP-3 cable, simultaneous operation of two 100BASE-T2 links, or one 100BASE-T2 and one 10BASE-T link, is permitted. Since voice-grade cables exhibit more signal attenuation and significantly higher crosstalk coupling between adjacent pairs than data-grade cables, sophisticated digital signal processing techniques are needed to achieve reliable 100 Mbit/s transmission. The 100BASE-T2 standard defines dual-duplex baseband transmission at a modulation rate of 25 MBaud. During each modulation interval, a four-bit data nibble or Ethernet specific control information is encoded into a pair of quinary signals. These signals are transmitted simultaneously on the two wire pairs in both signalling directions. In the receivers, adaptive digital filters are required for echo and NEXT cancellation, equalization, and interference suppression. Giovanni Cherubini, John Creigh, Sedat Ölçer, Sailesh K. Rao, Gottfried Ungerboeck |
ICC (2) | 4 |
| 1994 | Application Specific Memories for ATM Packet SwitchingabstractThe constrained data access patterns occurring within memory-based packet switches permit the design of application specific SRAM devices that may outperform generic SRAM parts in switch applications. We describe two such devices: one reads and writes a single location in a single 10 ns cycle; the other uses a systolic approach to pipeline accesses in a large array resulting in a 5 ns cycle time.> Alex G. Dickinson, Chris J. Nicol, Sailesh K. Rao, Mehdi Hatamian |
ISCAS | 3 |
| 1992 | The Rectilinear Steiner Arborescence Problem
Sailesh K. Rao, P. Sadayappan, Frank K. Hwang, Peter W. Shor |
Algorithmica | 1 |
| 1990 | A design environment for high performance VLSI signal processingabstractAn environment for the full-custom design of high-sample-rate digital signal processing (DSP) VLSI circuits is described. An overall design methodology that allows for tradeoffs between algorithms, architecture, and layout is presented. Key CAD tools used in this methodology include architecture mapping, generators, symbolic layout, clock network analysis, and timing simulation. These tools allow designers to move rapidly from specification to layout while keeping a close check on performance parameters such as speed and power dissipation. Two high-speed video chips that were designed using these techniques are described.> Sailesh K. Rao, Mehdi Hatamian, Bryan D. Ackland |
ICCD | 1 |
| 1990 | On the Analysis of Synchronous Computing SystemsabstractThis paper is concerned with the analysis of synchronous, special purpose, multiple-processor systems (including, e.g., systolic arrays). The analysis problem is that of determining the algorithm executed by the system. There has been some prior work in this area, especially by Melhem and Rheinboldt [SIAM J. Comput.,13 (1984), pp. 541–565], who were the first to obtain a general solution. The approach used here is different and apparently simpler. By combining ideas well known in system theory with certain graph-theoretical concepts, a simple procedure for recovering, within a natural equivalence, the iterative algorithm executed by a given special purpose synchronous computing array is obtained. The solution is based on reversing (modulo equivalence) the process by which an iterative algorithm is translated into a logical circuit. Juan-Manuel Jover, Thomas Kailath, Hanoch Lev-Ari, Sailesh K. Rao |
SIAM J. Comput. | 4 |
| 1989 | The matrix transform chipabstractThe matrix transform chip (MTC) is designed to perform matrix computations of the form Y=UDV where D is the input data matrix of 16-bit twos complement fixed-point numbers and U, V, are arbitrary coefficient matrices of the same precision. The data matrix D is input to the chip in raster scanned order at a maximum sample rate of 40 MHz, and the output matrix is provided in the same order. On a single chip, the maximum dimension of all matrices must be less than eight, but multiple chips can be cascaded to obtain arbitrary dimensions. The MTC consists of 16 16-bit parallel multipliers/40-bit accumulators, a kilobyte of dual-ported transposition static RAM, and a kilobyte of coefficient static RAM, arranged to interact in a regular iterative architecture. At peak operation, the MTC is capable of performing 0.64 billion fixed-point multiples, 0.64 billion 40-bit accumulates, along with 1.92 billion pseudorandom memory-access operations per second.> Sailesh K. Rao |
ICCD | 1 |
| 1989 | Communication reduction for distributed sparse matrix factorization on a processor meshabstractThe problem of reducing the amount of interprocessor communication during the distributed factorization of a sparse matrix on a mesh-connected processor network is investigated. Two strategies are evaluated - 1) use of a fragmented distribution of row/columns of the matrix to limit the number of processors to which each row/column segment is transmitted, and 2) use of the elimination tree to permute the matrix so as to internalize as much of the communication as possible. Empirical evaluation of the schemes using matrices derived from circuit simulation shows significant reduction in the amount of communication for a 64 processor mesh. P. Sadayappan, Sailesh K. Rao |
SC | 2 |
| 1988 | Regular iterative algorithms and their implementation on processor arraysabstractSome recent results are summarized concerning a class of algorithms known as regular iterative algorithms, particularly with respect to their implementations on processor arrays. Regular iterative algorithms contain all algorithms executed by systolic arrays as a proper subclass and are therefore of considerable importance in real-time signal processing applications. Some general concepts concerning the design of parallel architectures are introduced, and the importance of devising special techniques that utilize any available structure in the algorithm is highlighted. A generic description of the existing methodologies for the systematic design of systolic arrays is given. Using some simple examples, the limitations of these methods are shown. A formal methodology is proposed that overcomes the difficulties in the existing procedures.> Sailesh K. Rao, Thomas Kailath |
Proc. IEEE | 1 |
| 1987 | Array architectures for iterative algorithmsabstractRegular mesh-connected arrays are shown to be isomorphic to a class of so-called regular iterative algorithms. For a wide variety of problems it is shown how to obtain appropriate iterative algorithms and then how to translate these algorithms into arrays in a systematic fashion. Several "systolic" arrays presented in the literature are shown to be specific cases of the variety of architectures that can be derived by the techniques presented here. These include arrays for Fourier Transform, Matrix Multiplication, and Sorting. H. V. Jagadish, Sailesh K. Rao, Thomas Kailath |
Proc. IEEE | 2 |
| 1984 | Pipelined orthogonal digital lattice filtersabstractAn algorithm is presented for the design of systolic arrays that implement single-input single-output time-invariant digital filters. The algorithm is then specialized to the case where the realized array consists only of orthogonal rotational modules and delay elements interconnected in such a manner as to render the circuit pipelineable. Sailesh K. Rao, Thomas Kailath |
ICASSP | 1 |