Amit Rao

dblp:07/2220 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
code generation
0.011999
Storage Assignment Optimizations to Generate Compact and Efficient Code on Embedded DSPs · PLDI 1999
Compilers and program optimization › code generation › address code generation
offset assignment
0.011999
Storage Assignment Optimizations to Generate Compact and Efficient Code on Embedded DSPs · PLDI 1999
Compilers and program optimization
storage assignment
0.011999
Storage Assignment Optimizations to Generate Compact and Efficient Code on Embedded DSPs · PLDI 1999
Embedded and real-time systems › embedded software
DSP code generation
0.011999
Storage Assignment Optimizations to Generate Compact and Efficient Code on Embedded DSPs · PLDI 1999

Methods — techniques the papers use, named apart from their topics

algebraic transformations · 0.0algebraic transformation · 0.0
YearPublicationVenuePosition
2003 Optimal task scheduling at run time to exploit intra-tile parallelism
Fabrice Rastello, Amit Rao, Santosh Pande
Parallel Comput.2
1999 Storage Assignment Optimizations to Generate Compact and Efficient Code on Embedded DSPs
abstract
DSP architectures typically provide dedicated memory address generation units and indirect addressing modes with auto-increment and auto-decrement that subsume address arithmetic calculation. The heavy use of auto-increment and auto-decrement indirect addressing require DSP compilers to perform a careful placement of variables in storage to minimize address arithmetic instructions to generate compact and efficient DSP code. Liao et al. [11] formulated the problem of storage assignment as the simple o set assignment problem (SOA) and the general offset assignment problem (GOA), and proposed heuristic solutions. The storage allocation of variables critically depends on the sequence of variable accesses. In this paper we present techniques to optimize the access sequence of variables by applying algebraic transformations (such as commutativity and associativity) on expression trees to obtain the least cost offset assignment. We develop a new formulation of this problem as the least cost acces...
Amit Rao, Santosh Pande
PLDI1
1998 Optimal Task Scheduling to Minimize Inter-Tile Latencies
abstract
This work addresses the issue of exploiting intra-tile parallelism by overlapping communication with computation removing the restriction of atomicity of tiles. The effectiveness of tiling is then critically dependent on the execution order of tasks within a tile. We present a theoretical framework based on equivalence classes that provides an optimal task ordering under assumptions of constant and different permutations of tasks in individual tiles. Our framework is able to handle constant but compile-time unknown dependences by generating optimal task permutations at run-time and results in significantly lower loop completion times. Our solution is an improvement over previous approaches (Chou and Kung, 1993) (Dion et al., 1995) and is optimal for all problem instances. We also propose efficient algorithms that provide the optimal solution. The framework has been implemented as an optimization pass in the SUIF compiler and has been tested on a distributed memory system using a message passing model. We show that the performance improvement over previous results is substantial.
Fabrice Rastello, Amit Rao, Santosh Pande
ICPP2