EDBT 2026 Demo / reviewers in the wild / expert
Sanjay V. Rajopadhye
dblp:r/SanjayVRajopadhye
· DBLP profile ↗
70ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0002-4246-6066ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 57 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 11 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Maximal Simplification of Polyhedral ReductionsabstractReductions combine collections of input values with an associative and often commutative operator to produce collections of results. When the same input value contributes to multiple outputs, there is an opportunity to reuse partial results, enabling reduction simplification . Simplification often produces a program with lower asymptotic complexity. Typical compiler optimizations yield, at best, a constant fold speedup, but a complexity improvement from, say, cubic to quadratic complexity yields unbounded speedup for sufficiently large problems. It is well known that reductions in polyhedral programs may be simplified automatically , but previous methods cannot exploit all available reuse. This paper resolves this long-standing open problem, thereby attaining minimal asymptotic complexity in the simplified program. We propose extensions to prior work on simplification to support any independent commutative reduction. At the heart of our approach is piece-wise simplification, the notion that we can split an arbitrary reduction into pieces and then independently simplify each piece. However, the difficulty of using such piece-wise transformations is that they typically involve an infinite number of choices. We give constructive proofs to deal with this and select a finite number of pieces for simplification. Louis Narmour, Tomofumi Yuki, Sanjay V. Rajopadhye |
Proc. ACM Program. Lang. | 3 |
| 2024 | Taking RNA-RNA Interaction to Machine PeakabstractRNA-RNA interactions (RRIs) are essential in many biological processes, including gene transcription, translation, and localization. They play a critical role in diseases such as cancer and Alzheimer’s. Algorithms to model RRI typically use dynamic programming and have the complexity$\Theta (N^{3} \, M^{3})$in time and$\Theta (N^{2} \, M^{2})$in space where$N$and$M$are the lengths of the two RNA sequences. This makes it both essential and challenging to parallelize them. Previous efforts to do so have been hand-optimized, which is prone to human error and costly to develop and maintain. This paper presents a multi-core CPU parallelization of BPMax, one of the simpler RRI algorithms, generated by a user-guided polyhedral code generation tool,AlphaZ. The user starts with a mathematical specification of the dynamic programming algorithm and provides the choice of polyhedral program transformations such as schedules, memory-maps, and multi-level tiling.AlphaZautomatically generates highly optimized code. At the lowest level, we implemented a small hand-optimized register-tiled “matrix max-plus” kernel and integrated it with our tool-generated optimized code. Our final optimized program version is about$400\times$faster than the base program, translating to around 312 GFLOPS, more than half of our platform'sRoofline Machine Peak(RMP) performance. On a single core, we attain 80% ofRMP. The main kernel in the algorithm, whose complexity is$\Theta (N^{3} \, M^{3})$, attains 58 GFLOPS on a single-core and 344 GFLOPS on multi-core (90% and 58% ofRMP, respectively). Chiranjeb Mondal, Sanjay V. Rajopadhye |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Automatic Algorithm-Based Fault Tolerance (AABFT) of Stencil ComputationsabstractIn this work, we study fault tolerance of transient errors, such as those occurring due to cosmic radiation or hardware component aging and degradation, using Algorithm-Based Fault Tolerance (ABFT). ABFT methods typically work by adding some additional computation in the form of invariant checksums which, by definition, should not change as the program executes. By computing and monitoring checksums, it is possible to detect errors by observing differences in the checksum values. However, this is challenging for two key reasons: (1) it requires careful manual analysis of the input program, and (2) care must be taken to subsequently carry out the checksum computations efficiently enough for it to be worth it. Prior work has shown how to apply ABFT schemes with low overhead for a variety of input programs. Here, we focus on a subclass of programs called stencil applications, an important class of computations found widely in various scientific computing domains. We propose a new compilation scheme to automatically analyze and generate the checksum computations. To the best of our knowledge, this is the first work to do such a thing in a compiler. We show that low overhead code can be easily generated and provide a preliminary evaluation of the tradeoff between performance and effectiveness. Louis Narmour, Steven Derrien, Sanjay V. Rajopadhye |
PACT | 3 |
| 2023 | Distributed non-negative RESCAL with automatic model selection for exascale dataabstractWith the boom in the development of computer hardware and software, social media, IoT platforms, and communications, there has been exponential growth in the volume of data produced worldwide. Among these data, relational datasets are growing in popularity as they provide unique insights regarding the evolution of communities and their interactions. Relational datasets are naturally non-negative, sparse, and extra-large. Relational data usually contain triples (subject, relation, object) and are represented as graphs/multigraphs, called knowledge graphs, which need to be embedded into a low-dimensional dense vector space. Among various embedding models, RESCAL allows the learning of relational data to extract the posterior distributions over the latent variables and to make predictions of missing relations. However, RESCAL is computationally demanding and requires a fast and distributed implementation to analyze extra-large real-world datasets. Here we introduce a distributed non-negative RESCAL algorithm for heterogeneous CPU/GPU architectures with automatic selection of the number of latent communities (model selection), called pyDRESCALk. We demonstrate the correctness of pyDRESCALk with real-world and large synthetic tensors and the efficacy showing near-linear scaling that concurs with the theoretical complexities. Finally, pyDRESCALk determines the number of latent communities in an 11-terabyte dense and 9-exabyte sparse synthetic tensor. Manish Bhattarai, Namita Kharat, Ismael Boureima, Erik Skau, Ben Nebgen, Hristo N. Djidjev, Sanjay V. Rajopadhye, James P. Smith, Boian S. Alexandrov |
J. Parallel Distributed Comput. | 7 |
| 2023 | Increasing FPGA Accelerators Memory Bandwidth With a Burst-Friendly Memory LayoutabstractOffloading compute-intensive kernels to hardware accelerators relies on the large degree of parallelism offered by these platforms. However, the effective bandwidth of the memory interface often causes a bottleneck, hindering the accelerator’s effective performance. Techniques enabling data reuse, such as tiling, lower the pressure on memory traffic but do not fully exploit the bandwidth. A further increase in bandwidth utilization is possible by using burst rather than element-wise accesses, provided the data is contiguous in memory. In this article, we propose a memory allocation technique, and provide a proof-of-concept source-to-source compiler pass, that enables such burst transfers by modifying the data layout in external memory. Our experiments show the new memory allocation yields close to 100% bandwidth utilization while the memory engines occupy less than 5% of the field-programmable gate array (FPGA) logic area. Corentin Ferry, Tomofumi Yuki, Steven Derrien, Sanjay V. Rajopadhye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | BPPart: RNA-RNA Interaction Partition Function in the Absence of EntropyabstractA few classes of RNA-RNA interaction (RRI) with complex roles in cellular functions, such as miRNA-target and lncRNAs, have already been studied. Accordingly, RRI bioinformatics tools proposed in the last decade are tailored for those specific classes. Interestingly, there are somewhat unnoticed mRNA-mRNA interactions in the literature with potentially drastic biological roles. Hence, there is a need for high-throughput generic RRI bioinformatics tools that can be used in more comprehensive settings. In this work, we revisit two of the RRI partition function algorithms, piRNA and rip. These are equivalent methods that implement the most comprehensive and computationally intensive thermodynamic model for RRI. We propose simpler models that are shown to retain the vast majority of the thermodynamic information that the more complex models capture. Specifically, we simplify the energy model by ignoring the system’s entropy and show its equivalency to a base-pair counting model. We allow different weights for base-pairs to maximize the correlations with the full thermodynamic model. Our newly developed algorithm, BPPart, is 225× faster than piRNA and is more expressive and easier to analyze due to its simplicity and order of magnitude reduction in the number of dynamic programming tables. Still, based on our analysis of both the real and randomly generated data, its scores achieve a correlation of 0.855 with piRNA at 37^{∘}C. Finally, we illustrate one use-case of such simpler models to generate hypotheses about the roles of specific RNAs in various diseases. We have made our tool publicly available and believe that this faster and more expressive model will make the incorporation of physics-guided information in complex RRI analysis and prediction models more accessible. Ali Ebrahimpour Boroojeny, Sanjay V. Rajopadhye, Hamidreza Chitsaz |
WABI | 2 |
| 2020 | Revisiting Sparse Dynamic Programming for the 0/1 Knapsack ProblemabstractThe 0/1-Knapsack Problem is a classic NP-hard problem. There are two common approaches to obtain the exact solution: branch-and-bound (BB) and dynamic programming (DP). A so-called, “sparse” DP algorithm (SKPDP) that performs fewer operations than the standard algorithm (KPDP) is well known. To the best of our knowledge, there has been no quantitative analysis of the benefits of sparsity. We provide a careful empirical evaluation of SKPDP and observe that for a “large enough” capacity, C, the number of operations performed by SKPDP is invariant with respect to C for many problem instances. This leads to the possibility of an exponential improvement over the conventional KPDP. We experimentally explore SKPDP over a large range of knapsack problem instances and provide a detailed study of the attributes that impact the performance. Tarequl Islam Sifat, Nirmal Prajapati, Sanjay V. Rajopadhye |
ICPP | 3 |
| 2020 | LLOV: A Fast Static Data-Race Checker for OpenMP ProgramsabstractIn the era of Exascale computing, writing efficient parallel programs is indispensable, and, at the same time, writing sound parallel programs is very difficult. Specifying parallelism with frameworks such as OpenMP is relatively easy, but data races in these programs are an important source of bugs. In this article, we propose LLOV, a fast, lightweight, language agnostic, and static data race checker for OpenMP programs based on the LLVM compiler framework. We compare LLOV with other state-of-the-art data race checkers on a variety of well-established benchmarks. We show that the precision, accuracy, and the F1 score of LLOV is comparable to other checkers while being orders of magnitude faster. To the best of our knowledge, LLOV is the only tool among the state-of-the-art data race checkers that can verify a C/C++ or FORTRAN program to be data race free. Utpal Bora 0001, Pankaj Kukreja, Saurabh Joshi 0001, Ramakrishna Upadrasta, Sanjay V. Rajopadhye |
ACM Trans. Archit. Code Optim. | 6 |
| 2020 | Optimization Approach to Accelerator CodesignabstractWe propose an optimization approach for determining both hardware and software parameters for the efficient implementation of a (family of) applications called dense stencil computations on programmable general purpose computing on graphics processing units. We first introduce a simple, analytical model for the silicon area usage of accelerator architectures and a workload characterization of stencil computations. We combine this characterization with a parametric execution-time model and formulate a mathematical optimization problem that seeks to maximize a common objective function of all the hardware and software parameters. The solution to this problem, therefore, “solves” the codesign problem: simultaneously choosing software-hardware parameters to optimize total performance. We validate this approach by proposing architectural variants of the NVIDIA Maxwell GTX-980 (respectively, Titan X) specifically tuned to a predetermined workload of four common 2-D stencils (Heat, Jacobi, Laplacian, and Gradient) and two 3-D ones (Heat and Laplacian). Our model predicts that performance would potentially improve by 28% (respectively, 33%) with simple tweaks to the hardware parameters, such as tuning the number of streaming multiprocessors, the number of compute cores each contains, and the size of shared memory. We also develop a number of insights about the optimal regions of the design landscape. Nirmal Prajapati, Sanjay V. Rajopadhye, Hristo N. Djidjev, Nandakishore Santhi, Tobias Grosser, Rumen Andonov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | A Code Generator for Energy-Efficient Wavefront Parallelization of Uniform Dependence ComputationsabstractEnergy is now critical in all aspects of computing. We address a class of programs that includes so-called “stencil computations.” We address energy optimization of such programs. Since optimizing for speed alone already minimizes energy for most components, we seek to further improve the energy consumption by reducing the total number of off-chip memory accesses without sacrificing execution time. Our strategy uses two-level tiling: we first partition the iteration space into “passes,” each of which is tiled and parallelized. Here, the schedules that map the original program to the multi-pass, parametrically tiled code are specified by polynomials. They are more general than affine multidimensional schedules used by state of the art polyhedral compilers, so generating such codes automatically is an important open problem, and goes beyond the motivation of energy efficiency. We develop a parametric tiled code generator supporting our energy-efficient parallelization strategy. We give a simple linear regression model for energy as a function of performance counters. Our experimental validation on three platforms shows a reduction by about 74 percent (resp. 75 and 67 percent) of the dynamic memory energy consumption on an 8-core Xeon E5-2650 v2 (resp. 6-core Xeon E5-2620 v2 and 6-core Xeon E52602 v3). This leads to a reduction in the total energy of the program by 2 to 14 percent. Sanjay V. Rajopadhye |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | One size does not fit all: Implementation trade-offs for iterative stencil computations on FPGAsabstractIterative stencils are kernels in various application domains such as numerical simulations and medical imaging, that merit FPGA acceleration. The best architecture depends on many factors such as the target platform, off-chip memory bandwidth, problem size, and performance requirements. We generate a family of FPGA stencil accelerators targeting emerging System on Chip platforms, (e.g., Xilinx Zynq or Intel SoC). Our designs come with design knobs to explore trade-offs. We also propose performance models to hone in on the most interesting design points, and show how they accurately lead to optimal designs. The optimal choice depends on problem sizes and performance goals. Gaël Deest, Tomofumi Yuki, Sanjay V. Rajopadhye, Steven Derrien |
FPL | 3 |
| 2017 | Simple, Accurate, Analytical Time Modeling and Optimal Tile Size Selection for GPGPU StencilsabstractStencil computations are an important class of compute and data intensive programs that occur widely in scientific and engineeringapplications. A number of tools use sophisticated tiling, parallelization, and memory mapping strategies, and generate code that relies on vendor-supplied compilers. This code has a number of parameters, such as tile sizes, that are then tuned via empirical exploration. Nirmal Prajapati, Waruna Ranasinghe, Sanjay V. Rajopadhye, Rumen Andonov, Hristo N. Djidjev, Tobias Grosser |
PPoPP | 3 |
| 2015 | Automatic Energy Efficient Parallelization of Uniform Dependence ComputationsabstractEnergy is now a critical concern in all aspects of computing. We address a class of programs that includes the so-called "stencil computations" that have already been optimized for speed. We target the energy expended in dynamic memory accesses, since most other components of the total energy are usually already reduced when optimizing for speed alone. For a standard shared memory multi-core processor, we seek to minimize the total number of off-chip memory accesses without sacrificing execution time. Our strategy uses two-level tiling with multiple pipelined passes. Because of the sophisticated tiling and parallelization, such codes are difficult to write by hand, especially for parametric tile sizes. They are also beyond the capability of current code generators because the schedules used are polynomial functions, more general than multidimensional schedules. We implement a parametric tiled code generator to support this strategy, and also develop a simple quantitative linear regression model for the energy consumed by a program. We experimentally validate our techniques on a set of benchmarks including those from the Polybench suite on two platforms. Our experiments show that about 78% (resp. 80%) of the dynamic memory energy consumption on an 8-core Xeon E5-2650 v2 (resp. 6-core Xeon E5-2620 v2) based machine can be avoided. This leads to a reduction in the total energy of the program by 2% to 14%. Sanjay V. Rajopadhye |
ICS | 2 |
| 2014 | On Program Equivalence with Reductions
Guillaume Iooss, Christophe Alias, Sanjay V. Rajopadhye |
SAS | 3 |
| 2013 | Array dataflow analysis for polyhedral X10 programsabstractThis paper addresses the static analysis of an important class of X10 programs, namely those with finish/async parallelism, and affine loops and array reference structure as in the polyhedral model. For such programs our analysis can certify whenever a program is deterministic or flags races. Tomofumi Yuki, Paul Feautrier, Sanjay V. Rajopadhye, Vijay A. Saraswat |
PPoPP | 3 |
| 2012 | Scan detection and parallelization in "inherently sequential" nested loop programsabstractMost automatic parallelizers are based on detection of independent computations, and most of them cannot do anything if there is a true dependence between computations. However, this can be surmounted for programs that perform prefix computations (scans). We present a method for automatically parallelizing such "inherently sequential" programs. Our method, which handles arbitrarily nested loops, identifies situations where the computation performed by the loop body is equivalent to a matrix vector product over a semi-ring. We also deal with mutually dependent variables in the loop. Our method is implemented in a polyhedral program transformation and code generation system and generates OpenMP code. We also present strategies to improve the performance of the generated code, an analytical performance model for the expected speedup, as well as a method to choose the parallelization parameters optimally. We show experimentally that the scan parallelizations performed by our system are effective, yielding linear (iso-efficient) speedup in situations where no other parallelism is available. Sanjay V. Rajopadhye |
CGO | 2 |
| 2012 | Bridging the chasm between MDE and the world of compilation
Jean-Marc Jézéquel, Benoît Combemale, Steven Derrien, Clément Guy, Sanjay V. Rajopadhye |
Softw. Syst. Model. | 5 |
| 2012 | Parameterized loop tilingabstractLoop tiling is a widely used program optimization that improves data locality and enables coarse-grained parallelism. Parameterized tiled loops, where the tile sizes remain symbolic parameters until runtime, are quite useful for iterative compilers and autotuners that produce highly optimized libraries and codes. Although it is easy to generate such loops for (hyper-) rectangular iteration spaces tiled with (hyper-) rectangular tiles, many important computations do not fall into this restricted domain. In the past, parameterized tiled code generation for the general case of convex iteration spaces being tiled by (hyper-) rectangular tiles has been solved with bounding box approaches or with sophisticated and expensive machinery. We present a novel formulation of the parameterized tiled loop generation problem using a polyhedral set called the outset . By reducing the problem of parameterized tiled code generation to that of generating standard loops and simple postprocessing of these loops, the outset method achieves a code generation efficiency that is comparable to existing code generation techniques, including those for fixed tile sizes. We compare the performance of our technique with several other tiled loop generation methods on kernels from BLAS3 and scientific computations. The simplicity of our solution makes it well suited for use in production compilers—in particular, the IBM XL compiler uses the inset-based technique introduced in this article for register tiling. We also provide a complete coverage of parameterized tiling of perfect loop nests by describing three related techniques: (i) a scheme for separating full and partial tiles; (ii) a scheme for generating tiled loops directly from the abstract syntax tree representation of loops; (iii) a formal characterization of parameterized loop tiling using bilinear forms and a Symbolic Fourier-Motzkin Elimination (SFME)-based parameterized tiled loop generation method. Lakshminarayanan Renganarayanan, DaeGon Kim, Michelle Mills Strout, Sanjay V. Rajopadhye |
ACM Trans. Program. Lang. Syst. | 4 |
| 2011 | Model-Driven Engineering and Optimizing Compilers: A Bridge Too Far?
Antoine Floch, Tomofumi Yuki, Clément Guy, Steven Derrien, Benoît Combemale, Sanjay V. Rajopadhye, Robert B. France |
MoDELS | 6 |
| 2010 | Automatic creation of tile size selection modelsabstractTiling is a widely used loop transformation for exposing/exploiting parallelism and data locality. Effective use of tiling requires selection and tuning of the tile sizes. This is usually achieved by hand-crafting tile size selection (TSS) models that characterize the performance of the tiled program as a function of tile sizes. The best tile sizes are selected by either directly using the TSS model or by using the TSS model together with an empirical search. Hand-crafting accurate TSS models is hard, and adapting them to different architecture/compiler, or even keeping them up-to-date with respect to the evolution of a single compiler is often just as hard. Instead of hand-crafting TSS models, can we automatically learn or create them? In this paper, we show that for a specific class of programs fairly accurate TSS models can be automatically created by using a combination of simple program features, synthetic kernels, and standard machine learning techniques. The automatic TSS model generation scheme can also be directly used for adapting the model and/or keeping it up-to-date. We evaluate our scheme on six different architecture-compiler combinations (chosen from three different architectures and four different compilers). The models learned by our method have consistently shown near-optimal performance (within 5% of the optimal on average) across all architecture-compiler combinations. Tomofumi Yuki, Lakshminarayanan Renganarayanan, Sanjay V. Rajopadhye, Charles W. Anderson, Alexandre E. Eichenberger, Kevin O'Brien |
CGO | 3 |
| 2010 | Accelerating HMMER on FPGA using parallel prefixes and reductionsabstractHMMER is a widely used tool in bioinformatic, based on Profile Hidden Markov Models. The computation kernels of HMMER i.e. MSV and P7Viterbi are very compute intensive and data dependencies restrict to sequential execution. In this paper, we propose an original parallelization scheme for HMMER by rewriting their mathematical formulation, to expose the hidden potential parallelization opportunities. Our parallelization scheme targets FPGA technology, and our architecture can achieve 10 times speedup compared with that of latest HMMER3 SSE version, while not compromising on sensitivity of original algorithm. Naeem Abbas, Steven Derrien, Sanjay V. Rajopadhye, Patrice Quinton |
FPT | 3 |
| 2009 | A reindexing based approach towards mapping of DAG with affine schedules onto parallel embedded systems
Clémentin Tayou Djamégni, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset, Maurice Tchuenté |
J. Parallel Distributed Comput. | 3 |
| 2008 | A domain specific interconnect for reconfigurable computingabstractAffine Control Loops (ACLs) occur frequently in data- and computeintensive applications. Implementing ACLs directly on dedicated hardware has the potential for spectacular performance improvement in area, time and energy. An important challenge for such direct hardware compilation of ACLs is the interconnection between the different processing elements, which may be non-local as well as dynamic. We propose a generic, reconfigurable interconnection fabric which can realize the data-path of any ACL and be dynamically reconfigured in constant time. We have applied for a patent for this technology. Sanjay V. Rajopadhye, Gautam Gupta, Lakshminarayanan Renganarayanan |
LCTES | 1 |
| 2008 | Positivity, posynomials and tile size selectionabstractTiling is a widely used loop transformation for exposing/exploiting parallelism and data locality. Effective use of tiling requires selection and tuning of the tile sizes. This is usually achieved by developing cost models that characterize the performance of the tiled program as a function of tile sizes. All previous approaches to tile size selection (TSS) are cost model specific. Due to this they are neither extensible (e.g., to richer program classes/newer architectures) nor scalable (e.g., to multiple levels of tiling). This paper identifies positivity as a fundamental property shared by the functions and parameters commonly used in TSS models. We show how this positivity can be used as a basis to derive a TSS framework which is both efficient and scalable. We also show that almost all TSS models proposed in the literature (including those used in production compilers and auto-tuners) can be reduced to our framework. Lakshminarayanan Renganarayanan, Sanjay V. Rajopadhye |
SC | 2 |
| 2007 | 0/1 Knapsack on Hardware: A Complete SolutionabstractWe present a memory efficient, practical, systolic, parallel architecture for the complete 0/1 knapsack dynamic programming problem, including backtracking. This problem was intentionally selected because its dynamic dependencies introduce difficulties in hardware implementation. The architecture uses a divide-and-conquer technique that results in a pseudo-linear memory requirement. This memory reduction comes in exchange for a factor of two slowdown due to redundant computation. The architecture uses TΘ(𝓃 + 𝒑(С + 𝑾max)) memory and the run time is Θ(𝓃С/𝒑 + 𝓃log(𝓃/𝒑)). The heart of the architecture is a systolic module to compute the optimal profit for any problem that fits in available hardware resources. We implemented the module using 64 processors on an Alpha Data coprocessor board using a Xilinx VirtexII FPGA(2001 technology). Our implementation showed a factor of 32 improvement on the total execution time over a sequential algorithm running on a 1.5GHz Xeon processor(2000 technology) and a factor of 16 improvement over a 3.2GHz Pentium 4(2004 technology) and a 64 bit 3.4GHz Pentium 4(2006 technology). We measured complete wall-clock time, including the time to download the problem to the board, but not the time to download the bit stream to the FPGA. K. Nibbelink, Sanjay V. Rajopadhye, R. K. McConnell |
ASAP | 2 |
| 2007 | Scheduling in the Z-Polyhedral ModelabstractThe polyhedral model is extensively used for analyses and transformations of regular loop programs, one of the most important being automatic parallelization. The model, however, is limited in expressivity and the need for the generalization to more general class of programs has been widely known. Analyses and transformations in the polyhedral model rely on certain closure properties. Recently, these closure properties were extended to programs where variables may be defined over unions of Z-polyhedra which are the intersection of polyhedra and lattices. We present the scheduling analysis for the automatic parallelization of programs in the Z-polyhedral model, and obtain multidimensional schedules through an ILP formulation that minimizes latency. The resultant schedule can then be used to construct a space-time transformation to obtain an equivalent program in the Z-polyhedral model. Gautam Gupta, DaeGon Kim, Sanjay V. Rajopadhye |
IPDPS | 3 |
| 2007 | Towards Optimal Multi-level Tiling for Stencil ComputationsabstractStencil computations form the performance-critical core of many applications. Tiling and parallelization are two important optimizations to speed up stencil computations. Many tiling and parallelization strategies are applicable to a given stencil computation. The best strategy depends not only on the combination of the two techniques, but also on many parameters: tile and loop sizes in each dimension; computation-communication balance of the code; processor architecture; message startup costs; etc. The best choices can only be determined through design-space exploration, which is extremely tedious and error prone to do via exhaustive experimentation. We characterize the space of multi-level tilings and parallelizations for 2D/3D Gauss-Siedel stencil computation. A systematic exploration of a part of this space enabled us to derive a design which is up to a factor of two faster than the standard implementation. Lakshminarayanan Renganarayanan, Manjukumar Harthikote-Matha, Rinku Dewri, Sanjay V. Rajopadhye |
IPDPS | 4 |
| 2007 | Parameterized tiled loops for freeabstractParameterized tiled loops-where the tile sizes are not fixed at compile time, but remain symbolic parameters until later--are quite useful for iterative compilers and "auto-tuners" that produce highly optimized libraries and codes. Tile size parameterization could also enable optimizations such as register tiling to become dynamic optimizations. Although it is easy to generate such loops for (hyper) rectangular iteration spaces tiled with (hyper) rectangular tiles, many important computations do not fall into this restricted domain. Parameterized tile code generation for the general case of convex iteration spaces being tiled by (hyper) rectangular tiles has in the past been solved with bounding box approaches or symbolic Fourier Motzkin approaches. However, both approaches have less than ideal code generation efficiency and resulting code quality. We present the theoretical foundations, implementation, and experimental validation of a simple, unified technique for generating parameterized tiled code. Our code generation efficiency is comparable to all existing code generation techniques including those for fixed tile sizes, and the resulting code is as efficient as, if not more than, all previous techniques. Thus the technique provides parameterized tiled loops for free! Our "one-size-fits-all" solution, which is available as open source software can be adapted for use in production compilers. Lakshminarayanan Renganarayanan, DaeGon Kim, Sanjay V. Rajopadhye, Michelle Mills Strout |
PLDI | 3 |
| 2007 | The Z-polyhedral modelabstractThe polyhedral model is a well developed formalism and has been extensively used in a variety of contexts viz. the automatic parallelization of loop programs, program verification, locality, hardware generationand more recently, in the automatic reduction of asymptotic program complexity. Such analyses and transformations rely on certain closure properties. However, the model is limited in expressivity and the need for a more general class of programs is widely known. Gautam Gupta, Sanjay V. Rajopadhye |
PPoPP | 2 |
| 2007 | Multi-level tiling: M for the price of oneabstractTiling is a widely used loop transformation for exposing/exploiting parallelism and data locality. High-performance implementations use multiple levels of tiling to exploit the hierarchy of parallelism and cache/register locality. Efficient generation of multi-level tiled code is essential for effective use of multi-level tiling. Parameterized tiled code, where tile sizes are not fixed but left as symbolic parameters can enable several dynamic and run-time optimizations. Previous solutions to multi-level tiled loop generation are limited to the case where tile sizes are fixed at compile time. We present an algorithm that can generate multi-level parameterized tiled loops at the same cost as generating single-level tiled loops. The efficiency of our method is demonstrated on several benchmarks. We also present a method-useful in register tiling-for separating partial and full tiles at any arbitrary level of tiling. The code generator we have implemented is available as an open source tool. DaeGon Kim, Lakshminarayanan Renganarayanan, Dave Rostron, Sanjay V. Rajopadhye, Michelle Mills Strout |
SC | 4 |
| 2006 | An Improved Systolic Architecture for LU DecompositionabstractLU-Decomposition is a classic problem for which many systolic array implementations have been proposed, the best of which takes 3n-3 time on n2/2 PEs, for a dense, nxn matrix. In this paper, we first give a proof that if only nearest neighbor communication is allowed, this time is a lower bound. We then generalize it to 2n + n/k - 3 time if one allows k-bounded broadcasts (i.e., if it takes m/k time steps to broadcast a value to m destination nodes). We also present a new architecture with this improved execution time, which uses n2/2 PEs, each one consisting of two multiplier-subtractor units, but active only on alternate cycles. This leads to a speedup and efficiency of kn2/(6k+3) and 2k/(6k + 3) respectively. For k = 1, our proposed architecture achieves the performance of the best previously known systolic array implementation. Special cases of our results include similar improvements to algorithms for solving (upper and lower) triangular linear systems by (forward and backward) substitution. DaeGon Kim, Sanjay V. Rajopadhye |
ASAP | 2 |
| 2006 | Simplifying reductionsabstractWe present optimization techniques for high level equational programs that are generalizations of affine control loops (ACLs). Significant parts of the SpecFP and PerfectClub benchmarks are ACLs. They often contain reductions: associative and commutative operators applied to a collection of values. They also often exhibit reuse: intermediate values computed or used at different index points being identical. We develop various techniques to automatically exploit reuse to simplify the computational complexity of evaluating reductions. Finally, we present an algorithm for the optimal application of such simplifications resulting in an equivalent specification with minimum complexity. Gautam Gupta, Sanjay V. Rajopadhye |
POPL | 2 |
| 2004 | A Geometric Programming Framework for Optimal Multi-Level TilingabstractDetermining the optimal tile size-one that minimizes the execution time-is a classical problem in compilation and performance tuning of loop kernels. Designing a model of the overall execution time of a tiled loop nest is an important subproblem. Both problems become harder when tiling is applied at multiple levels. We present a framework for determining the optimal tile sizes for a fully permutable, perfectly nested, rectangular loop with uniform dependences. Our framework supports multiple levels of tiling and uses a BSP style high level model for estimating the overall execution time of a loop program. In our framework, the problem of determining the optimal tile sizes, subject to memory capacity and bandwidth constraints, is modeled as a geometric program and transformed into a convex optimization problem, which can be solved efficiently. The model is validated through experimental results obtained by running twenty loop programs for different levels of tiling and different program and tile parameters. Our framework is very general and can also be used to solve the optimal tile size problem with many other models of execution time. Lakshminarayanan Renganarayanan, Sanjay V. Rajopadhye |
SC | 2 |
| 2003 | Switched Memory Architectures-Moving Beyond Systolic ArraysabstractAlthough current ASIC, FPGA and reconfigurable computing technologies support on-chip memories and hardware reconfiguration, these features are not exploited by systolic arrays and their associated synthesis methods. We propose a new architectural model called switched memory architecture (SMA) to overcome these limitations. SMAS are (strictly) more powerful than systolic arrays, are suitable for a wide range of target technologies, and can be derived through the well developed design methodology of the polyhedral model. We illustrate the power of SMAs by showing how any SARE with a one dimensional schedule can be implemented as an SMA without any slowdown. We formally characterize the class of allocation functions that are suitable for SMAs and also describe a systematic procedure for deriving SMAs from SAREs. Lakshminarayanan Renganarayanan, Sanjay V. Rajopadhye |
ASAP | 2 |
| 2003 | Optimal Semi-Oblique TilingabstractFor 2D iteration space tiling, we address the problem of determining the tile parameters that minimize the total execution time on a parallel machine. We consider uniform dependency computations tiled so that (at least) one of the tile boundaries is parallel to the domain boundaries. We determine the optimal tile size as a closed form solution. In addition, we determine the optimal number of processors and also the optimal slope of the oblique tile boundary. Our results are based on the BSP model, which assures the portability of the results. Our predictions are justified on a sequence global alignment problem specialized to similar sequences using Fickett's k-band algorithm, for which our optimal semi-oblique tiling yields an improvement of a factor of 2.5 over orthogonal tiling. Our optimal solution requires a block-cyclic distribution of tiles to processors. The best one can obtain with only block distribution (as many authors require) is three times slower. Furthermore, our best running time is within 10 percent of the "predicted theoretical peak" performance of the machine!. Rumen Andonov, Stephan Balev, Sanjay V. Rajopadhye, Nicola Yanev |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2002 | Scheduling reductions on realistic machinesabstractMany computations can be modeled with systems of affine recurrence equations (SAREs) over polyhedral domains. We study the problem of scheduling individual computations of an SARE in the presence of reductions i.e., operations specifying the accumulation of a set of values to produce a single value. Reductions involve a commutative and associative operator and therefore, per se, do not impose any specific order. However, on realistic machines, operators have bounded fan-in and therefore an order of accumulation (serialization) is needed. Arbitrary serializations may adversely affect the running time of a program. We develop an algorithm to determine efficient serializations of all reductions. We illustrate our methods with two significant examples. Gautam Gupta, Sanjay V. Rajopadhye, Patrice Quinton |
SPAA | 2 |
| 2001 | Combining Instruction and Loop Level Parallelism for FPGAs
Steven Derrien, Sanjay V. Rajopadhye, Susmita Sur-Kolay |
FCCM | 2 |
| 2001 | Loop Tiling for Reconfigurable Accelerators
Steven Derrien, Sanjay V. Rajopadhye |
FPL | 2 |
| 2001 | Uniformization of Affine Dependance Programs for Parallel Embedded System DesignabstractThe paper is concerned with the uniformization of a system of affine recurrence equations. This transformation is used in the design (or compilation) of highly parallel embedded systems (VLSI systolic arrays, signal processing filters, etc.). We present and implement an automatic system to achieve uniformization of systems of affine recurrence equations. We unify the results from many earlier papers, develop some theoretical extensions, and then propose effective uniformization algorithms. Our results can be used in any high level synthesis tool based on polyhedral representation of nested loop computations. Manju Manjunathaiah, Graham M. Megson, Sanjay V. Rajopadhye, Tanguy Risset |
ICPP | 3 |
| 2001 | Proving Properties of Multidimensional Recurrences with Application to Regular Parallel AlgorithmsabstractWe present a set of verification methods to prove properties of parallel systems described by means of multidimensional affine recurrence equations. We use polyhedral analysis and transformation techniques together with theorem proving. Polyhedral techniques allow us to handle simple but otherwise costly proof steps, while theorem proving provides more expressivity and more complex proof techniques. This allows large, generic and structured systems to be verified. These methods are implemented in the M-MAlpha environment using the PVS theorem prover. 1 David Cachera, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset |
IPDPS | 3 |
| 2001 | Optimal semi-oblique tilingabstractFor 2-D iteration space tiling, we address the problem of determining the tile parameters that minimize the total execution time under the BSP model. We consider uniform dependency computations, tiled so that (at least) one of the tile boundaries is parallel to the domain boundary. We determine the optimal tile size as a closed form solution. In addition, we determine the optimal number of processors and also the optimal slope of the oblique tile boundary. Rumen Andonov, Stephan Balev, Sanjay V. Rajopadhye, Nicola Yanev |
SPAA | 3 |
| 2000 | Quadratic Control Signals in Linear Systolic ArraysabstractWe describe a new problem in the control of systolic arrays that arises with multidimensional time called the quadratic control problem. We present a solution for the problem as well as an implementation. The solution is based on difference calculus and provides an elegant and scalable method for handling control signals in quadratic time. Scott Bowden, Doran Wilde, Sanjay V. Rajopadhye |
ASAP | 3 |
| 2000 | FCCMS and the Memory WallabstractAlthough there has been considerable work in the conventional general purpose processors community on how to tackle an important looming problem, we are not aware of any similar effort for custom computing machines. The aim of this paper is to analyze the state of the art, pose the relevant questions, and indicate a preliminary solution vis a vis the following question: how will custom computing machines face the memory wall. Steven Derrien, Sanjay V. Rajopadhye |
FCCM | 2 |
| 2000 | Derivation of systolic algorithms for the algebraic path problem by recurrence transformations
Clémentin Tayou Djamégni, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset |
Parallel Comput. | 3 |
| 2000 | Optimizing memory usage in the polyhedral modelabstractThe polyhedral model provides a single unified foundation for systolic array synthesis and automatic parallelization of loop programs. We investigate the problem of memory reuse when compiling Alpha (a functional language based on this model). Direct compilation would require unacceptably large memory (for example O(n 3 ) for matrix multiplication). Researchers have previously addressed the problem of memory reuse, and the analysis that this entails for projective memory allocations. This paper addresses, for a given schedule, the choice of the projections so as to minimize the volume of the residual memory. We prove tight bounds on the number of linearly independent projection vectors. Our method is constructive, yielding an optimal memory allocation. We extend the method to modular functions, and deal with the subsequent problems of code generation. Our ideas are illustrated on a number of examples generated by the current version of the Alpha compiler. Fabien Quilleré, Sanjay V. Rajopadhye |
ACM Trans. Program. Lang. Syst. | 2 |
| 1999 | The Algebraic Path Problem Revisited
Sanjay V. Rajopadhye, Claude Tadonki, Tanguy Risset |
Euro-Par | 1 |
| 1998 | Optimal Orthogonal Tiling
Rumen Andonov, Sanjay V. Rajopadhye, Nicola Yanev |
Euro-Par | 2 |
| 1998 | Linear Programming Models for Scheduling Systems of Affine Recurrence Equations - A Comparative Study
Stephan Balev, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset |
SPAA | 3 |
| 1997 | Optimal Orthogonal Tiling of 2-D Iterations
Rumen Andonov, Sanjay V. Rajopadhye |
J. Parallel Distributed Comput. | 2 |
| 1997 | Multirate VLSI Arrays and Their SynthesisabstractMany applications in signal and image processing can be efficiently implemented on regular VLSI architectures such as systolic arrays. Multirate arrays (MRAs) are an extension of systolic arrays where different data streams are propagated with different clocks. We address the analysis and synthesis problem for this class of architectures. We present a formal definition of MRAs, as systems of recurrence equations defined over sparse polyhedral domains. We also give transformation rules for this class of recurrences, and use them to show that MRAs constitute a particular subset of systems of affine recurrence equations (SoAREs). We then address the synthesis problem, and show how an MRA can be systematically derived from an initial specification in the form of a mathematical equation. The main transformations that we use are domain rescalings and dependency decomposition, and we illustrate our method by deriving a hitherto unknown decimation filter array. Patrick M. Lenders, Sanjay V. Rajopadhye |
IEEE Trans. Computers | 2 |
| 1997 | Knapsack on VLSI: from Algorithm to Optimal CircuitabstractWe present a parallel solution to the unbounded knapsack problem on a linear systolic array. It achieves optimal speedup for this well-known, NP-hard problem on a model of computation that is weaker than the PRAM. Our array is correct by construction, as it is formally derived by transforming a recurrence equation specifying the algorithm. This recurrence has dynamic dependencies, a property that puts it beyond the scope of previous methods for automatic systolic synthesis. Our derivation thus serves as a case study. We generalize the technique and propose a systematic method for deriving systolic arrays by nonlinear transformations of recurrences. We give sufficient conditions that the transformations must satisfy, thus extending systolic synthesis methods. We address a number of pragmatic considerations: implementing the array on only a fixed number of PEs, simplifying the control to just two counters and a few latches, and loading the coefficients so that successive problems can be pipelined without any loss of throughput. Using a register level model of VLSI, we formulate a nonlinear optimization problem to minimize the expected running time of the array. The analytical solution of this problem allows us to choose the memory size of each PE in an optimal manner. Rumen Andonov, Sanjay V. Rajopadhye |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1996 | Extension Of The Alpha Language To Recurrences On Sparse Periodic DomainsabstractALPHA is a functional language based on systems of affine recurrence equations over polyhedral domains. We present an extension of ALPHA to deal with sparse polyhedral domains. Such domains are modeled by Z-polyhedra, namely the intersection of lattices and polyhedra. We summarize the mathematical closure properties of Z-polyhedra, and we show how the important features of ALPHA, namely normalization, substitution, change of basis, are preserved in the extension. Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset |
ASAP | 2 |
| 1996 | Two-dimensional orthogonal tiling: from theory to practiceabstractIn pipelined parallel computations the inner loops are often implemented in a block fashion. In such programs, an important compiler optimization involves the need to statically determine the grain size. This paper presents extensions and experimental validation of the previous results of Andonov and Rajopadhye (1994) on optimal grain size determination. Rumen Andonov, Hafid Bourzoufi, Sanjay V. Rajopadhye |
HiPC | 3 |
| 1996 | Parallel Divide and Conquer on MeshesabstractWe address the problem of mapping divide-and-conquer programs to mesh connected multicomputers with wormhole or store-and-forward routing. We propose the binomial tree as an efficient model of parallel divide-and-conquer and present two mappings of the binomial tree to the 2D mesh. Our mappings exploit regularity in the communication structure of the divide-and-conquer computation and are also sensitive to the underlying flow control scheme of the target architecture. We evaluate these mappings using new metrics which are extensions of the classical notions of dilation and contention. We introduce the notion of communication slowdown as a measure of the total communication overhead incurred by a parallel computation. We conclude that significant performance gains can be realized when the mapping is sensitive to the flow control scheme of the target architecture. Virginia Mary Lo, Sanjay V. Rajopadhye, Jan Arne Telle, Xiaoxiong Zhong |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1995 | Synthesis of Multirate VLSI ArraysabstractMany applications in signal and image processing can be implemented on regular VLSI architectures such as systolic arrays. Multirate arrays, or MRAs are an extension of systolic arrays where different data streams propagate with different clocks. It is known that they can be modelled as systems of uniform recurrence equations over sparse polyhedral domains. Using well known linear index transformation rules for systems of affine recurrence equations, or SAREs, we show that MRAs constitute a particular proper subset of SAREs. We describe how an MRA can be systematically derived from an initial specification in the form of a mathematical equation. The main transformation that we use is dependency decomposition, and rue illustrate our method by deriving a hitherto unknown decimation filter array that improves upon the hardware cost of previously published filters. Patrick M. Lenders, Sanjay V. Rajopadhye |
ASAP | 2 |
| 1995 | The naive execution of affine recurrence equationsabstractIn recognition of the fundamental relation between regular arrays and systems of affine recurrence equations, the ALPHA language was developed as the basis of a computer aided design methodology for regular array architectures. ALPHA is used to initially specify algorithms at a very high algorithmic level. Regular array architectures can then be derived from the algorithmic specification using a transformational approach supported by the ALPHA environment. This design methodology guarantees the final design to be correct by construction, assuming the initial algorithm was correct. In this paper, we address the problem of validating an initial specification. We demonstrate a translation methodology which compiles ALPHA into the imperative sequential language C. The C-code may then be compiled and executed to test the specification. We show how an ALPHA program can be naively implemented by viewing it as a set of monolithic arrays and their filing functions, implemented using applicative caching. This is the approach which is used by the translator. We discuss two problems that had to be solved before implementing the translator. The first is how to allocate 1-dimensional storage for a polyhedron, and the second is how to scan a polyhedron with nested loops. Doran Wilde, Sanjay V. Rajopadhye |
ASAP | 2 |
| 1994 | A sparse knapsack algo-tech-cuit and its synthesisabstractWe systematically derive an improved algorithm (called the sparse algorithm) for the general knapsack problem which has better average case performance than the standard (dense) dynamic programming algorithm. The derivation is based on transformation of the standard recurrences into stream functional programs, and cannot be achieved by the usual space-time mapping techniques because the dependencies are statically unpredictable. Furthermore such a sparse algorithm for the general knapsack problem has not been proposed in the literature, to the best of our knowledge. We also implement the sparse algorithm on a linear asynchronous array with constant size memory on each PE (i.e., a wavefront array processor). Using LPGS partitioning, the algorithm can run on an arbitrary size ring and has optimal time speedup.> Rumen Andonov, Sanjay V. Rajopadhye |
ASAP | 2 |
| 1993 | An optimal algo-tech-cuit for the knapsack problemabstractThe authors first present a formal derivation and proof of correctness of a systolic array for the knapsack problem, an NP-complete problem whose dependency graph is not completely known statically. With q PEs, each with a fixed size memory, the arraystretch runs in /spl Gamma/(mc/q), which gives optimal speedup of the algorithm. However, it has an intricate tag-based control mechanism which is difficult to implement, and the authors improve the architecture so that the control can be implemented with two simple counters and a few flip-flops. Cofficient loading is done with a multi-rate clock which avoids the need for shadow registers. The authors then explore the tradeoff between the number of PEs and memory size, /spl alpha/, using the expected running time of the algorithm as a cost measure and a register level model of VLSI. It is shown analytically how /spl alpha/ may be chosen to optimize the total computation time, yielding an area time optimal circuit.> Rumen Andonov, Sanjay V. Rajopadhye |
ASAP | 2 |
| 1993 | An improved systolic algorithm for the algebraic path problem
Sanjay V. Rajopadhye |
Integr. | 1 |
| 1991 | A folding transformation for VLSI IIR filter array designabstractAn array structure of infinite impulse response (IIR) digital filters is presented. The architectures are formally derived using techniques for synthesizing systolic arrays from high-level (algorithmic) specifications. First, the authors apply the conventional synthesis techniques for deriving the array structure for IIR filters and then they discuss a folding transformation. Folding is a technique used for reducing the array size by allocating several computations onto the same processing element. It enables the designer to explore additional design options which are not available in the conventional approach of using linear projections to derive systolic arrays.> Sanjay V. Rajopadhye, Sayfe Kiaei |
ICASSP | 1 |
| 1991 | Synthesizing fully efficient systolic arraysabstractIt is shown that in arrays derived by the conventional integral linear transformation, there always exists a basic direction nu such that one and only one out of delta consecutive processors is active at any time unit along any line parallel to nu . Therefore, one can merge delta neighboring processors along lines parallel to nu and derive a new array which is fully efficient. The new array has only 1/ delta processors and has the same computation time complexity. The method is constructive: once the user has chosen the timing function and allocation function, the required nu can be generated automatically. After choosing nu , the appropriate quasi-affine allocation function can also be automatically derived, as can the additional registers, wires, and control for the processors. This indicates that it is always possible to synthesize a fully efficient systolic array in a practical system.> Xiaoxiong Zhong, Sanjay V. Rajopadhye |
ICASSP | 2 |
| 1990 | Scheduling affine parameterized recurrences by means of Variable Dependent Timing FunctionsabstractThe authors present new scheduling techniques for systems of affine recurrence equations. They show that it is possible to extend earlier results on affine scheduling to the case when each variable of the system is scheduled independently of the others by an affine timing-function. This new technique makes it possible to analyze systems of recurrence equations with variables in different index spaces, and multi-step systolic algorithms. This theory applies directly to many problems, such as dynamic programming, LU decomposition, and 2-D convolution, and it avoids in particular preliminary heuristic rewriting of the equations.> Christophe Mauras, Patrice Quinton, Sanjay V. Rajopadhye, Yannick Saouter |
ASAP | 3 |
| 1990 | OREGAMI: Software Tools for Mapping Parallel Computations to Parallel Architectures
Virginia Mary Lo, Sanjay V. Rajopadhye, Samik Gupta, David Keldsen, Moataz A. Mohamed, Jan Arne Telle |
ICPP (2) | 2 |
| 1990 | Mapping Divide-and-Conquer Algorithms to Parallel Architectures
Virginia Mary Lo, Sanjay V. Rajopadhye, Samik Gupta, David Keldsen, Moataz A. Mohamed, Jan Arne Telle |
ICPP (3) | 2 |
| 1990 | Automating the design of systolic arrays
Sanjay V. Rajopadhye, Richard M. Fujimoto |
Integr. | 1 |
| 1990 | Synthesizing systolic arrays from recurrence equations
Sanjay V. Rajopadhye, Richard M. Fujimoto |
Parallel Comput. | 1 |
| 1989 | Synthesizing Systolic Arrays with Control Signals from Recurrence Equations
Sanjay V. Rajopadhye |
Distributed Comput. | 1 |
| 1986 | On Synthesizing Systolic Arrays from Recurrence Equations with Linear Dependencies
Sanjay V. Rajopadhye, S. Purushothaman Iyer, Richard M. Fujimoto |
FSTTCS | 1 |
| 1986 | Verification of Systolic Arrays: A Stream Function Approach
Sanjay V. Rajopadhye, Prakash Panangaden |
ICPP | 1 |
| 1985 | Formal semantics for a symbolic IC design technique: Examples and applications
Sanjay V. Rajopadhye, P. A. Subrahmanyam |
Integr. | 1 |