Ranga Vemuri

dblp:81/2940 · DBLP profile ↗
← Back
135ranked-venue papers
6as first author
4since 2021 · last 2022
0000-0002-4903-2746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 128 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 30 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorTheory of computation · 4Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2022 Efficient Method for Timing-based Information Flow Verification in Hardware Designs
abstract
Timing side channels are a serious threat to the security of hardware designs. By analyzing the execution times of a design, the attacker can expose the secret information. This paper proposes an approach to verify and monitor timing-based information flow properties. In addition, the method can highlight the path that is vulnerable to leakage, making it easier to trace the leaking channel. The method can be used during formal verification, dynamic verification during simulation, post-fabrication validation, and run-time monitoring if one is necessary. The method reduces the overhead of the security model, which helps speed up the verification process and create an efficient run-time hardware monitor. Various timing-based information flow properties from five different hardware designs were verified. The results show that our approach can accurately detect hardware timing channels with lower overhead.
Khitam Alatoun, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2022 Model Checking Leveraged Error Localization for Complex RTL Designs
abstract
Function and security verification has emerged as an important concern in the design of systems-on-chip (SoC) architectures. Bug detection and localization traces the cause of any violation to a compact portion of the design. In this paper, given a design’s correctness and security properties, we present a novel methodology to localize errors in complex RTL models. Also, we discuss how our approach is applied to a large SoC design to debug its security policy violations. We have leveraged the features of a model checker and assertion based verification to achieve this task. To show a variety of debug scenarios, we inject design mutations in the OpenRISC-1200 SoC causing security policy violations in the form of assertion failures, yielding counterexample traces. Through these traces, our algorithm automatically identifies a concise list of signals responsible for causing the security violation.
Suriya Srinivasan, Ranga Vemuri
ICCD2
2022 FPGA-Based Stochastic Local Search Satisfiability Solvers Exploiting High Bandwidth Memory
abstract
Boolean Satisfiability (SAT) problems for realistic applications are becoming increasingly large and complex, making it difficult for deterministic methods to be used on modern CPUs. In this paper we present hardware Stochastic Local Search (SLS) solvers that utilize a state-of-the-art FPGA-based accelerator. The method uses well-known SLS heuristics implemented on an accelerator with large memory capacity using the Vitis HLS system. The solver can be reconfigured with multiple SLS kernels to take advantage of the resources on the Alveo U280 accelerator. Combined, our techniques achieve faster convergence on large SAT problems compared to CPU solvers.
Christopher Chuvalas, Ranga Vemuri
VLSI-SoC2
2021 Efficient Methods for SoC Trust Validation Using Information Flow Verification
abstract
Information flow properties are essential to identify security vulnerabilities in System-on-Chip (SoC) designs. Verifying information flow properties, such as integrity and confidentiality, is challenging as these properties cannot be handled using traditional assertion-based verification techniques. This paper proposes two novel approaches, a universal method and a property-driven method, to verify and monitor information flow properties. Both methods can be used for formal verification, dynamic verification during simulation, post-fabrication validation, and run-time monitoring. The universal method expedites implementing the information flow model and has less complexity than the most recently published technique. The property-driven method reduces the overhead of the security model, which helps speed up the verification process and create an efficient run-time hardware monitor. More than 20 information flow properties from 5 different designs were verified and several bugs were identified. We show that the method is scalable for large systems by applying it to an SoC design based on an OpenRISC-1200 processor.
Khitam Alatoun, Shanmukha Murali Achyutha, Ranga Vemuri
ICCD3
2019 A State Machine Encoding Methodology Against Power Analysis Attacks
Richa Agrawal, Ranga Vemuri, Mike Borowczak
J. Electron. Test.2
2018 Improving the Security of Split Manufacturing Using a Novel BEOL Signal Selection Method
abstract
Split manufacturing of integrated circuits (IC) was proposed as a possible defense against security issues arising from the use of potentially untrusted foundries. However, split manufactured designs were shown to be vulnerable to a new form of attack known as the proximity attack which attempts to reverse engineer the BEOL (Back End of Line) signals. Hence, care must be exercised in identifying the BEOL signals and their placement and routing. In this paper, we present a secure BEOL signal selection algorithm to defeat proximity attacks. Our method is based on two novel features: First, we introduce a new metric for signal selection based on the effect each signal has on the outputs. Second, we use a multiway partitioning algorithm to find a 'secure' cut-set which is the set of signals assigned to the BEOL layers. Our approach increases the number of BEOL nets while minimizing the impact on performance. We present experimental results which show significant improvement in security (on average by 800%) with only a modest effect on performance (less than 4% on average).
Suyuan Chen, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2018 Reverse Engineering of Split Manufactured Sequential Circuits Using Satisfiability Checking
abstract
Split Manufacturing was proposed as a promising strategy to thwart reverse engineering and Trojan insertion at untrusted foundries. However, attack methods based on physical design hints have been proposed to reverse engineer a combinational circuit with Front-End-Of-line (FEOL) layers only. But none of them can guarantee 100% recovery of BEOL signals since no validation can be done during the attack process. In this paper, we introduce an attack flow that can recover 100% BEOL signals for sequential circuits effectively. Our approach shows promising results in attacking sequential circuits without access to flip-flop outputs. We demonstrate the effectiveness of the new attack method on a set of sequential benchmarks from ISCAS-89 and ITC-99 sets which have been widely used in related research. Results show logic equivalence between the original circuit and the recovered circuit for all benchmarks.
Suyuan Chen, Ranga Vemuri
ICCD2
2018 On the Effectiveness of the Satisfiability Attack on Split Manufactured Circuits
abstract
Split manufacturing of integrated circuits was proposed as a strong defense technique against reverse engineering at untrusted foundries. Split manufacturing has been shown to be vulnerable to proximity-based attacks and suitable defenses against these attacks have been proposed. In this paper, we apply an attack against split manufactured circuits based on satisfiability (SAT) solving, without requiring any proximity information. Our method formulates the problem of recovering the hidden signals as a Boolean decryption problem of determining the control signals (keys) of a multiplexer network, determines logical constraints to avoid solutions which could enable cyclic paths through the circuit and reduces the number of these constraints by eliminating redundant constraints. The resulting constrained decryption problem is solved by a satisfiability attack. We demonstrate the effectiveness of the attack on a set of benchmarks. In addition, we discuss a split manufacturing approach to potentially deter the SAT attack.
Suyuan Chen, Ranga Vemuri
VLSI-SoC2
2017 Effective Signal Restoration in Post-Silicon Validation
abstract
Trace buffer based run-time techniques have been proposed and widely applied to post-silicon validation for bug detection and localization after post-manufacturing testing. However, hardware overhead limits the size of the trace buffer, i.e. the number of signals traced during run-time. In order to alleviate this limitation, maximal state restoration based on the data stored in trace buffer is desirable with minimal compromise in efficiency. Forward propagation and backward justification (FB) methods were commonly employed and satisfiability-based (SAT) techniques were recently introduced for signal restoration. In this paper, we propose efficient heuristics based on the structure of the circuit and a signal ordering technique for signal restoration using hybrid FB and SAT methods. Experimental results show that our reconvergent fanout based algorithm can achieve up to 72% improvement and our signal ordering heuristic can attain 77.9% improvement compared to the state-of-the-art methods while maintaining optimal restoration ratio.
Xiaobang Liu, Ranga Vemuri
ICCD2
2016 A novel simulation based approach for trace signal selection in silicon debug
abstract
With the fabrication technology fast approaching 7nm, post-silicon validation has become an integral part of integrated circuit design to capture and eliminate functional bugs that escape pre-silicon validation. The major roadblock in post-silicon functional verification is limited observability of internal signals in a design. A possible solution to address this roadblock is to make use of embedded memories on chip called trace buffers. The amount of debug data that can be acquired from the trace buffer depends on its width and depth. The width of the trace buffer limits the number of signals that can be traced and the depth of the trace buffer limits the number of samples that can be acquired. Using the acquired data from the trace buffer, the values of other nodes in the circuit can be reconstructed. These trace buffers have limited area, hence only a few critical signals can be recorded by it. In this work we used the simulated annealing heuristic to select trace signals. We developed this idea from the fact that trace signal selection can be viewed as a bi-partitioning problem, the set of flip-flops being tapped onto the trace buffer is one partition and remaining flip-flops form the other partition. Experimental results demonstrate that our approach can result in better restoration ratio compared to the state-of-the-art techniques.
Prabanjan Komari, Ranga Vemuri
ICCD2
2016 Fast Inversions in Small Finite Fields by Using Binary Trees
abstract
Inversions in small finite fields are playing a key role in many areas. We present techniques to exploit binary trees for fast inversions in |$GF(2^n)$| and |$GF(p)$| , where |$n$| is a positive integer and |$p$| is a prime number. The non-pipelined versions of our design in |$GF(2^n)$| and |$GF(p)$| have the execution time of |$(n-1)(T_{AND}+T_{XOR})$| and |$\lfloor \log _2p\rfloor (T_{AND}+T_{XOR})$| , where |$T_{AND}$| and |${T_{XOR}}$| are delays of AND and XOR gates, respectively. The pipelined version of our design has a throughput rate of one result per |$T_{AND}$| (or |$T_{XOR}$| ). The latency is the greater value between |$T_{AND}$| and |$T_{XOR}$| . In other words, the time complexities of non-pipelined and pipelined versions are |$O(n)$| (or |$O(log_2p)$| ) and |$O(1)$| , respectively. Experimental results and comparisons show that our design provides significant reductions in both the execution time and time–area product, e.g. the execution time of inversion in |$GF(2^{12})$| is reduced by 73 |$\%$| and time–area product of inversion in |$GF(2^6)$| is reduced by 77 |$\%$| .
Haibo Yi, Shaohua Tang, Ranga Vemuri
Comput. J.3
2011 Aggressive Runtime Leakage Control Through Adaptive Light-Weight Vth Hopping With Temperature and Process Variation
abstract
The increasing leakage power consumption and stringent thermal constraint necessitate more aggressive leakage control techniques. Power gating and body biasing are widely used for standby leakage control. Their large energy overhead for performing mode transition is the major obstacle for more aggressive leakage control. Temperature and process variation (TV/PV) further magnify the overhead problem, leading to so-called “corner case leakage control” problem. Light-weightVthhopping (LW-VH) is a candidate technique to tackle the energy overhead problem. This paper demonstrates the application of LW-VH on microarchitectural- and RTL-level idleness exploitation with adaptive control techniques for TV/PV compensation. Adaptive LW-VH shows 30% average saving on total CPU leakage at microarchitectural level, and 4% to 15% leakage saving at RTL level. By combining all the techniques proposed in this paper, a three-tier aggressive leakage control system is introduced to fully exploit idleness at all levels.
Hao Xu 0010, Wen-Ben Jone, Ranga Vemuri
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Dynamic Characteristics of Power Gating During Mode Transition
abstract
With the technology moving into the deep sub-100-nm region, the increase of leakage power consumption necessitates more aggressive power reduction techniques. Power gating is a promising technique. Our research emphasizes that with the latest and future technologies, power gating operates frequently in its transition mode, especially for aggressive leakage reduction. The dynamic characteristics of power gating during its mode transition is critical for making design decision. Hence we derive a fast, accurate, and temperature-aware model to characterize the dynamic behavior of power gating during mode transition. The applications of this model include the estimation of several key design parameters for power gating, such as dynamic virtual ground voltage, dynamic leakage variation and energy break-even time. It provides an efficient estimation engine for power gating design optimization. The accuracy of the model has been verified by extensive HSPICE experiments. The model is computationally efficient due to the usage of various approximation methods.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Stretching the limit of microarchitectural level leakage control with Adaptive Light-Weight Vth Hopping
abstract
Power gating (PG) and body biasing (BB) are popular leakage control techniques at microarchitectural level. However, their large overhead prevents them from being applied for active leakage reduction. The overhead problem is further magnified by temperature and process variation, leading to the “corner case leakage control” problem. This paper presents an Adaptive Light-Weight Vth Hopping technique. This technique dramatically reduces the overhead for mode transition, addresses the corner case leakage control problem, and thus enables active leakage control.
Hao Xu 0010, Wen-Ben Jone, Ranga Vemuri
ICCAD3
2010 Current shaping and multi-thread activation for fast and reliable power mode transition in multicore designs
abstract
Power gating has been widely adopted in multicore designs. The design of fast and reliable power mode transition for per-core power gating remains a challenging problem. This paper studies the design methodology for fast power gating wake-up with guaranteed power integrity. Two novel techniques, namely current shaping and multi-thread activation are proposed. Models and physical implementation of both techniques are analyzed. Experimental results demonstrated 1.5 to 11 times wake-up time speedup with no penalty on area or power consumptions by using the proposed techniques.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
ICCAD2
2009 A graph grammar based approach to automated multi-objective analog circuit design
abstract
This paper introduces a graph grammar based approach to automated topology synthesis of analog circuits. A grammar is developed to generate circuits through production rules, that are encoded in the form of a derivation tree. The synthesis has been sped up by using dynamically obtained design-suitable building blocks. Our technique has certain advantages when compared to other tree-based approaches like GP based structure generation. Experiments conducted on an opamp and a vco design show that unlike previous works, we are capable of generating both manual-like designs (bookish circuits) as well as novel designs (unfamiliar circuits) for multi-objective analog circuit design benchmarks.
Angan Das, Ranga Vemuri
DATE2
2009 Selective light Vth hopping (SLITH): Bridging the gap between runtime dynamic and leakage
abstract
Ever since the invention of various leakage power reduction techniques, leakage and dynamic power reduction techniques are categorized into two separate sets. Most of them cannot be applied together during runtime. The gap between them is due to the large energy breakeven time (EBT) and wakeup time (WUT) of conventional leakage reduction techniques. This paper proposes a new leakage reduction technique (SLITH) based on Vthhopping. SLITH has very low EBT and WUT, yet keeps the effectiveness of leakage reduction. Thus, it is able to reduce the gap, and enables joint dynamic and leakage power reduction. SLITH can be applied together with clock gating, precomputation and operand isolation etc., and significantly reduces both dynamic and active leakage power consumption.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
DATE2
2009 A methodology for application-specific NoC architecture generation in a dynamic task structure environment
abstract
To avoid bandwidth violations, the NoC architecture generation phase must consider the bandwidth variation along various links. In this paper, we analyze the impact of the dynamic nature of task graphs on bandwidth requirements and present an algorithm to find the Minimum BandWidth Guarantee along the various links of the NoC architecture.
Balasubramanian Sethuraman, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2009 Temporal and spatial idleness exploitation for optimal-grained leakage control
abstract
Runtime leakage control techniques, such as power gating (PG) and body biasing (BB), have been applied in a coarse-grained manner traditionally. In order to enable more aggressive leakage reduction, researchers are seeking ways to control leakage with finer granularity. Our research proposes two novel methods, namely circuit clustering for temporal and spatial idleness exploitation, to systematically reduce the granularity of leakage control and improve leakage reduction. Another strength of this paper is the quantitative study of leakage saving and control cost by leakage control with different granularity. With our quantitative study, designers can make the trade-off between leakage saving and control cost, and decide the optimum granularity for leakage control. A heuristic algorithm has been developed to automate the two circuit clustering methods and determine the optimum granularity for any given circuit. The analysis and experiments of this paper is mainly based on RBB. They are also applicable to PG by modifying the cost function.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
ICCAD2
2009 Accurate estimation of vector dependent leakage power in the presence of process variations
abstract
With the increasing importance of run-time leakage power dissipation (around 55% of total power), it has become necessary to accurately estimate it not only as a function of input vectors but also as a function of process parameters. Leakage power corresponding to the maximum vector presents itself as a higher bound for run-time leakage and is a measure of reliability. In this work, we address the problem of accurately estimating the probabilistic distribution of the maximum runtime leakage power in the presence of variations in process parameters such as threshold voltage, critical dimensions and doping concentration. Both sub-threshold and gate leakage current are considered. A heuristic approach is proposed to determine the vector that causes the maximum leakage power under the influence of random process variations. This vector is then used to estimate the lognormal distribution of the total leakage current of the circuit by summing up the lognormal leakage current distributions of the individual standard cells at their respective input levels. The proposed method has been effective in accurately estimating the leakage mean, standard deviation and probability density function (PDF) of ISCAS-85 benchmark circuits. The average errors of our method compared with near exhaustive random vector testing for mean and standard deviation are 1.32% and 1.41% respectively.
Romana Fernandes, Ranga Vemuri
ICCD2
2008 Topology synthesis of analog circuits based on adaptively generated building blocks
abstract
This paper presents an automated analog synthesis tool for topology generation and subsequent circuit sizing. Though sizing is indispensable, the paper mainly concentrates on topology generation. A new kind of GA is developed, where a fraction of the offsprings in each generation is built from building blocks or cells obtained from previous generations. The cells are stored in a hierarchically arranged library that also contains information on the preferred neighborhood of each cell. The adaptively formed cell library starts only with basic elements and gradually includes functionally useful and bigger blocks, pertinent to the design. The techniques have been applied to synthesize an operational amplifier and a ring oscillator design. Results show that with reasonable computational effort, topologies have evolved that are designer understandable.
Angan Das, Ranga Vemuri
DAC2
2008 Fast Analog Circuit Synthesis Using Sensitivity Based Near Neighbor Searches
abstract
We present an efficient analog synthesis algorithm employing regression models of circuit matrices. Circuit matrix models achieve accurate and speedy synthesis of analog circuits. In this paper, synthesis is accelerated by eliminating numerous computations of the matrix elements during a synthesis run. Computations are avoided by reusing exact or nearby design points visited during previous synthesis iterations. Hashing and multidimensional nearest neighbor lookup are used in incremental evaluation of design solutions encountered during synthesis. Sensitivity of the design variables is considered for locating a neighboring solution. Neighbor lookup is efficiently performed using box-decomposition trees. The proposed method is used to synthesize three benchmark circuits. Results show that with hashing and neighbor lookup, synthesis is 6x-13x faster than with the use of matrix models alone.
Almitra Pradhan, Ranga Vemuri
DATE2
2008 A layout-aware analog synthesis procedure inclusive of dynamic module geometry selection
abstract
We propose an algorithm for sizing analog circuits using parasitic aware circuit matrix models. A novel scheme of separating schematic and parasitic models is proposed. As layout details are not abstracted in the circuit performance, the developed models can be used for different module geometries. Regression models developed make parasitic estimation much faster than a layout inclusive approach. The proposed approach is successfully used for dynamic module geometry selection during synthesis. Experiments conducted on operational amplifier and filter topologies demonstrate the accuracy of our proposed approach. For both circuits, results are within a mean error of 1 percent compared to an exact layout and spice approach.
Almitra Pradhan, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2008 Accurate energy breakeven time estimation for run-time power gating
abstract
Run-time Power Gating (RTPG) is a recent technique, which aims at aggressively reducing leakage power consumption. Energy breakeven time (EBT), or equivalent sleep time has been proposed as a critical figure of merit of RTPG. Our research introduces the definition of average EBT in a run-time environment. We develop a method to estimate the average EBT for any given circuit block, considering the impact of circuit states. HSPICE simulation results on ISCAS85 benchmark circuits show that the average EBT model has on the average 1.8% error. The CAD tool implemented based on the model can perform fast estimations with a speedup of 3000times over HSPICE.
Hao Xu 0010, Wen-Ben Jone, Ranga Vemuri
ICCAD3
2008 Run-time Active Leakage Reduction by power gating and reverse body biasing: An eNERGY vIEW
abstract
Run-time active leakage reduction (RALR) is a recent technique and aims at aggressively reducing leakage power consumption. This paper studies the feasibility of RALR from the energy aspect, for both power gating (PG) and reverse body bias (RBB) implementations.We develop two energy saving models for PG and RBB, respectively. These models can accurately estimate the circuit energy saving at any time, even when the circuit is in state transition. In PG modeling, we discover a physical phenomenon called ldquoinstant savingrdquo, which can affect the model accuracy by 30%-50%. Based on the RBB model, we derive the optimum design point of RBB for RALR. Finally in terms of energy saving, we define four figures-of-merit, to compare the efficacy of using PG and RBB to implement RALR.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
ICCD2
2008 ATLAS: An adaptively formed hierarchical cell library based analog synthesis framework
abstract
This paper presents ATLAS - a framework for automated analog circuit synthesis that comprises of both topology generation and subsequent circuit sizing. A hierarchically arranged building block or cell library is used in this regard. The adaptively formed library starts only with basic elements and gradually includes functionally useful and bigger blocks, pertinent to the design under consideration. The sizer is based on the simulated annealing algorithm, and HSPICE is used for performance evaluation. The tool has been used to synthesize an operational amplifier and a ring oscillator. Results show that with reasonable computational effort, designs and associated cells have evolved that are human understandable and comparable to hand-crafted designs.
Angan Das, Ranga Vemuri
ISCAS2
2008 Dynamic virtual ground voltage estimation for power gating
abstract
With the technology moving into the deep sub-100nm region, the increase of leakage power consumption necessitates more aggressive power reduction techniques. Power gating is a promising technique. Our research emphasizes the virtual ground voltage (VVG) as the key to make critical design trade-offs for power gating. We develop an accurate model to estimate the dynamic VVG value of a circuit block as a function of time after its ground is gated. Experimental results show that the model has less than 1% average error compared with HSPICE results. The CAD tool implemented based on the model has a 100 times speedup over HSPICE.
Hao Xu 0010, Ranga Vemuri, Wen-Ben Jone
ISLPED2
2007 Power variations of multi-port routers in an application-specific NoC design : A case study
abstract
In this research, we analyze the power variations present in a router having varied number of ports, in a Networks- on-Chip. The work is divided into two sections, projecting the merits and shortcomings of a multi-port router from the aspect of power consumption. First, we evaluate the power variations present during the transfers between various port pairs in a multi-port router. The power gains achieved through careful port selection during the mapping phase of the NoC design are shown. Secondly, through exhaustive experimentation, we discuss the IR-drop related issues that arise when using large multi-port routers.
Balasubramanian Sethuraman, Ranga Vemuri
ICCD2
2007 A spline based regression technique on interval valued noisy data
abstract
In this paper we present a spline based center and range method (SCRM) to perform regression on interval valued noisy data. The method provides a fast and accurate mechanism to model and predict upper and lower limits of unknown functions in a bounded design space. This technique is superior to previously existing techniques like center and range linear least square regression (CRM). The accurate models may find wide usage in high precision applications. The effectiveness of the proposed technique is demonstrated through experiments on datasets with various applications.
Balaji Kommineni, Shubhankar Basu, Ranga Vemuri
ICMLA3
2007 GAPSYS: A GA-based Tool for Automated Passive Analog Circuit Synthesis
abstract
This paper presents GAPSYS - a genetic algorithm based automated circuit synthesis tool for passive analog circuits. It describes the procedure for developing both the circuit topology and the component values for a passive analog circuit comprising of R, L and C components from a given set of specifications. The novelty of the work pertains to the component value assignment procedure for the initial set of circuits and the crossover techniques employed. Experiments conducted on two low-pass filter design benchmarks demonstrate the effectiveness of GAPSYS as a synthesis tool.
Angan Das, Ranga Vemuri
ISCAS2
2007 Multicasting based topology generation and core mapping for a power efficient networks-on-chip
abstract
Networks-on-Chip (NoC) is an emerging alternative for system integration that is projected to meet the growing communication demands for future System-on-Chips. Compared to the bus-based systems, traditional NoCs do not have versatile data transfer capabilities like broadcasting. Multi2 Router is a Multi Local Port Router (MLPR) architecture that has multicast feature in-built inside the router elements of an MLPR-based NoC. In this research,we present an NoC configuration generation approach exploiting the multicast feature. Compared to the traditional single port based unicast transfers, we observe an average of 50% packet reduction (maximum of 74% using 9 Local Port (LP) router, in benchmark p3), across a set of benchmarks. On an average, when compared to the traditional 1 LP unicast router, there is a 16% reduction in the execution time and 35% reduction (maximum of 67% in benchmark p4) in total power consumption. The results show the promise of the proposed scheme, and thus, help to realize power-efficient Networks-on-Chip.
Balasubramanian Sethuraman, Ranga Vemuri
ISLPED2
2007 Regression based circuit matrix models for accurate performance estimation of analog circuits
abstract
Automated analog circuit synthesis techniques depend on fast and reliable estimation of circuit performance. This paper presents a highly accurate method of estimating performances by constructing models of the circuit matrix instead of the traditionally used performance models. Device matching in analog circuits is utilized to identify identical elements in the circuit matrix and reduce the number of elements to be modeled. Experiments conducted on three operational amplifier topologies demonstrate the effectiveness of the method in achieving correct performance prediction. Results show that the performances can be predicted within a mean error of 0.1% compared to a SPICE simulation.
Almitra Pradhan, Ranga Vemuri
VLSI-SoC2
2007 Power invariant secure IC design methodology using reduced complementary dynamic and differential logic
abstract
Security of cryptographic devices (secure ICs) like smart cards has come under threat from powerful side channel attacks like Differential Power Analysis (DPA). DPA uses power consumption information leaked from the secure IC in conjunction with statistical correlation techniques to retrieve the secret key stored in the secure IC. The most effective countermeasure to resist DPA attacks is to make the power consumption of the secure IC invariant, hence uncorrelated to the input data (secret key). In hardware implementations, this can be achieved by designing the secure IC using Dynamic and Differential Logic (DDL) style. In this paper, we present a novel methodology to design DPA-resistant power invariant secure ICs using Reduced Complementary Dynamic and Differential Logic (RCDDL). The proposed methodology involves strategies to design: 1) RCDDL gates, and 2) secure circuits using RCDDL gates. Experiments show significant improvements in security strength, average power consumption and area, when compared with a similar secure DDL and non-secure static-CMOS logic design styles.
Vijay Sundaresan, Srividhya Rammohan, Ranga Vemuri
VLSI-SoC3
2006 optiMap: a tool for automated generation of noc architectures using multi-port routers for FPGAs
abstract
Networks-on-chip (NoC) way of system design has been introduced to overcome the communication and the performance bottlenecks of a bus based system design. Area is at a premium in FPGAs. In this research, we propose to reduce network area overhead by reducing the number of routers, by making the router handle multiple logic cores. We implement an improved multi-local port router design with variable number of local ports. In addition to substantial area savings, we observe significant performance improvement. We discuss the issues involved in the use of multi-local port routers for NoC design in FPGAs. We observe an average of 36% area savings (maximum of 47.5%) on XC2VP30 FPGA and significant performance gain (30% average compared to single-local port version) with a multi-local port router. Mapping of cores onto such a non-traditional NoC architecture is a complex task. We present an algorithm which optimally maps the cores based on the given set of objectives. For the given task graph and the set of constraints, the algorithm finds the optimal number of routers, configuration of each router, optimal mesh topology and the final mapping. We test the algorithm on a wide variety of benchmarks and report the results
Balasubramanian Sethuraman, Ranga Vemuri
DATE2
2006 Efficient temperature-dependent symbolic sensitivity analysis and symbolic performance evaluation in analog circuit synthesis
abstract
We present a new methodology for fast analog circuit synthesis, based on the use of temperature-dependent symbolic sensitivity analysis and symbolic performance evaluation in synthesis loop. Fast sensitivity analysis achieved and performance estimation are based on element-coefficient diagrams(ECDs). Sensitivity and performance evaluation expressions are generated from ECDs at the same time which reduces overall runtime greatly. The experimental results demonstrate that the speed and convergence of analog synthesis are improved significantly.
Huiying Yang, Ranga Vemuri
DATE2
2006 Multi2 Router: A Novel Multi Local Port Router Architecture with Broadcast Facility for FPGA-Based Networks-on-Chip
abstract
Modern FPGAs provide increased gate count with decreased power consumption. Several IP cores along with embedded processor and memory provide a great opportunity of implementing system-on-chip (SoC) designs on configurable devices. Networks-on-Chip (NoC) is an emerging style of SoC design, introduced to overcome the communication and performance bottlenecks of a shared-bus approach. Multi local port router (MLPR) present a novel design alternative for the traditional NoC design. This new methodology offers numerous advantages including bandwidth optimization and reduced network area & power consumption, resulting eventually in improved performance of the NoC system. Unlike the bus-based systems, communication in NoCs until now have been between pair of cores, with no scope of multi-casting. In this research, we advance a step further in the pursuit of a high performance FPGA-based NoC system. We exploit the multi-casting nature present in various application system task graphs and present a novel & improved MLPR architecture with broadcast capability. We present the modified architecture, the decoding scheme and the stripped-down crosspoint matrix, resulting in reduced logic usage & increased performance. We report the synthesis and the simulation results.
Balasubramanian Sethuraman, Ranga Vemuri
FPL2
2006 Transformation synthesis for data intensive applications to FPGAs
abstract
Without the adequate awareness of trade-off between different resources, it is extremely difficult for system synthesis tools to achieve high performance solutions when mapping the applications to FPGA-based computing engines. In this paper, we present an automatic synthesis methodology which attacks both memory and logic assignments by interacting with behavioral synthesis. The problem is formulated as part of the heuristic algorithm by exploiting application specific information and organizing possible data structures and computations for data-intensive applications. We have evaluated the proposed framework on a set of DSP benchmarks and a real multimedia application by generating register-transfer level (RTL) implementations. The results show that, by using our proposed techniques, it is possible the synthesized designs obtain significant (avg. of 34.8%) performance improvements over the conventional synthesis approaches.
Renqiu Huang, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2006 Studying a GALS FPGA architecture using a parameterized automatic design flow
abstract
Routing delays dominate other delays in current FPGA designs. We have proposed a novel Globally Asynchronous Locally Synchronous (GALS) FPGA architecture called the GAPLA to deal with this problem. In the GAPLA architecture, The FPGA area is divided into locally synchronous blocks and the communications between them are through asynchronous I/O interfaces. An automatic design flow is developed for the GAPLA architecture. Starting from behavioral description, a design is partitioned into smaller modules and fit to GAPLA synchronous blocks. The asynchronous communications between modules are then sytthesized. The CAD flow is parameterized in modeling the GAPLA architecture. By manipulating the parameters, we could study different factors of the designed GAPLA arcitecturc. Our experimental results show an average of 20% performance improvement could be achieved by the GAPLA architecture.
Ranga Vemuri
ICCAD2
2006 Exact hierarchical symbolic analysis of large analog networks using a general interconnection template
abstract
The primary focus of this paper is the development of a hierarchical symbolic analysis method, which can be used to generate symbolic performance models (SPMs) for large parasitic-inclusive analog circuits. In this paper, a new exact hierarchical technique is proposed, where transfer functions (TF) are synthesized for a general interconnection template (GIT) of two subcircuits. Extremely efficient element-coefficient diagrams (ECD) are used for symbolic analysis of subcircuits. The results of TF-synthesis of a GIT, lead to the development of an easily automatable symbolic analysis method. The efficiency and scope of this method is then demonstrated on few common large analog networks
Mukesh Ranjan, Ranga Vemuri
ISCAS2
2006 Hierarchical constraint transformation based on genetic optimization for analog system synthesis
Nagu R. Dhanwada, Alex Doboli, Adrián Núñez-Aldana, Ranga Vemuri
Integr.4
2006 Energy management for battery-powered reconfigurable computing platforms
abstract
We define portable reconfigurable computing platforms as those which have some form of configurable logic coupled with other on-chip or off-chip processing units such as soft processors, embedded processors, and voltage-scalable processors. In the first part of this paper, we present and test a unique methodology where we dynamically change the active area of a field programmable gate array (FPGA) to vary the battery usage and lifetime of the system, by running it on several different taskgraph structures and report an average of 14% and as high as 21%, less battery capacity used, as compared to nonoptimal execution. In the second part of this paper, we integrate the above methodology with more traditional voltage and frequency scaling techniques for portable systems and present a heuristic iterative algorithm for single and multiple processing units. The iterative heuristic algorithm finds a sequence of tasks along with an appropriate design point (implementation option) for each task, such that a deadline is met and the amount of battery energy used is as small as possible. We have used several real-world benchmarks to test the effectiveness of this methodology and we will present the results.
Jawad Khan, Ranga Vemuri
IEEE Trans. Very Large Scale Integr. Syst.2
2005 An error-driven adaptive grid refinement algorithm for automatic generation of analog circuit performance macromodels
abstract
In this paper, we present an error-driven adaptive sampling algorithm called adaptive grid refinement (AGR) algorithm to automatically generate performance macromodels for analog circuits. Starting from samples on a coarse grid, the AGR algorithm builds a global model and validates its accuracy on an independent validation data set sampled within this grid. If this model is not accurate enough on the validation data, the grid is split into equal sized smaller grids. On each of these grids, a local model is built using samples on this grid and its neighboring and validated similarly. A grid will not be further refined only if the corresponding local model is accurate on its validation data set. The algorithm will stop when all the local models are accurate on their corresponding validation data set. We build six performance macromodels of a CMOS opamp using the AGR algorithm and compare it with the competing techniques. The strengths and weaknesses of the proposed algorithm are discussed.
Mengmeng Ding, Glenn Wolfe, Ranga Vemuri
ASP-DAC3
2005 Using GALS architecture to reduce the impact of long wire delay on FPGA performance
abstract
Interconnect delay is becoming a major roadblock to FPGA performance with technology scaling and growing chip sizes. globally asynchronous locally synchronous (GALS) design is considered a potential solution to this issue. An important design decision in building a GALS FPGA architecture is to determine the appropriate GALS island size. A large GALS island will reduce the asynchronous communication overhead but the interconnect delay inside an island is increased. On the other hand, asynchronous communication overhead could be a major concern for a small GALS island size. In this paper, we propose a design flow to investigate this tradeoff. The input circuit is first divided into partitions according to the specified GALS island size and each partition is then implemented with commercially available CAD tools. The overall system performance is estimated by a performance evaluator. Experimental results validate our design flow and show a performance improvement of around 20% by adopting a GALS architecture.
Ranga Vemuri
ASP-DAC2
2005 Efficient symbolic sensitivity analysis of analog circuits using element-coefficient diagrams
abstract
This paper presents a new method to perform efficient first-order symbolic sensitivity analysis of analog circuits by direct differentiation of symbolic expressions stored as element-coefficient diagrams (ECDs). An ECD is a compact graphical representation of a symbolic transfer function. It is the cancellation-free and per-coefficient term generation version of determinant decision diagrams (DDDs). The symbolic sensitivity equations obtained from ECDs are stored as a sensitivity-ECDs(SECDs) and can be evaluated extremely fast as it inherits the properties of ECDs. The proposed methodology has been applied to the calculation of sensitivities of four benchmark circuits and it has been demonstrated to be as accurate and more efficient than numerical sensitivity analysis done by SPECTRE.
Huiying Yang, Mukesh Ranjan, Wim Verhaegen, Mengmeng Ding, Ranga Vemuri, Georges Gielen
ASP-DAC5
2005 A combined feasibility and performance macromodel for analog circuits
abstract
The need to reuse the performance macromodels of an analog circuit topology challenges existing regression based modeling techniques. A model of good reusability should have a number of independent design parameters and each parameter can vary in a large numeric range. On the other hand, these requirements can cause a large percentage of functionally incorrect designs in the design space and thus results in a sparse feasible design space. They also complicate the mathematical relationship between the performance parameters and the design parameters. In order to tackle these challenges, this paper presents a combined feasibility and performance macromodel based on Support Vector Machines (SVMs). The feasibility model identifies the feasible designs that satisfy the design constraints. The performance macromodel is valid for feasible designs. Feasibility macromodeling is formulated as a classification problem while performance macromeling as a regression problem. An active learning scheme [5] has been applied to improve the accuracy of the feasibility model much faster than only using uniformly distributed designs in the entire design space. Our experiment shows that the performance macromodels in the feasible design space are more accurate and faster to construct and evaluate than performance macromodels in the entire design space without functional or performance constraints considered.
Mengmeng Ding, Ranga Vemuri
DAC2
2005 Multi-Placement Structures for Fast and Optimized Placement in Analog Circuit Synthesis
abstract
The paper presents the novel idea of multi-placement structures, for a fast and optimized placement instantiation in analog circuit synthesis. These structures need to be generated only once for a specific circuit topology. When used in synthesis, these pre-generated structures instantiate various layout floorplans for the various sizes and parameters of a circuit. Unlike procedural layout generators, they enable fast placement of circuits while keeping the quality of the placements at a high level during the synthesis process. The fast placement is a result of high speed instantiation resulting from the efficiency of the multi-placement structure. The good quality of placements derives from the extensive and intelligent search process that is used to build the multi-placement structure. The target benchmarks of these structures are analog circuits in the vicinity of 25 modules. An algorithm for the generation of such multi-placement structures is presented. Experimental results show placement execution times with an average of a few milliseconds making them usable during layout-aware synthesis for optimized placements.
Raoul F. Badaoui, Ranga Vemuri
DATE2
2005 Inductive and Capacitive Coupling Aware Routing Methodology Driven by a Higher Order RLCK Moment Metric
abstract
A new routing methodology, which accounts for inductive and capacitive coupling between neighboring wires is proposed. The inductive and capacitive coupling of the wires are introduced through a 'moment' based higher order RLCK cost function. The routing process guided by this cost function ensures that the final solution has minimum ringing and delay.
Amitava Bhaduri, Ranga Vemuri
DATE2
2005 A Two-Level Modeling Approach to Analog Circuit Performance Macromodeling
abstract
We present a two-level modeling approach to performance macromodeling based on a radial basis function support vector machine (SVM). The two-level model consists of a feasibility model and a set of performance models. The feasibility model identifies the feasible designs that satisfy the design constraints. The performance macromodel is valid for feasible designs. We formulate the feasibility macromodeling problem as a classification problem and the performance macromodeling as a regression problem and apply an SVM algorithm to build the classifier and regressors correspondingly. Our experiment shows that performance macromodels for feasible designs are much more accurate, faster to train and evaluate than those without functional or performance constraints considered.
Mengmeng Ding, Ranga Vemuri
DATE2
2005 An Iterative Algorithm for Battery-Aware Task Scheduling on Portable Computing Platforms
abstract
We consider battery powered portable systems which either have field programmable gate arrays (FPGA) or voltage and frequency scalable processors as their main processing element. An application is modeled in the form of a precedence task graph at a coarse level of granularity. We assume that, for each task in the task graph, several unique design-points are available which correspond to different hardware implementations for FPGAs and different voltage-frequency combinations for processors. It is assumed that performance and total power consumption estimates for each design-point are available for any given portable platform, including the power usage of peripheral components, such as memory and display. We present an iterative heuristic algorithm which finds a sequence of tasks along with an appropriate design-point for each task, such that a deadline is met and the amount of battery energy used is as small as possible. A detailed illustrative example, along with a case study of a real-world application of a robotic arm controller which demonstrates the usefulness of our algorithm, is also presented.
Jawad Khan, Ranga Vemuri
DATE2
2005 The GAPLA: A Globally Asynchronous Locally Synchronous FPGA Architecture
abstract
This paper proposes GAPLA: a globally asynchronous locally synchronous programmable logic array architecture. The whole FPGA area is divided into locally synchronous blocks wrapped with asynchronous I/O interfaces. Data communications between synchronous blocks are controlled by 2-phase handshaking signals. The size and shape of each locally synchronous block are programmable so that different modules in a design can be effectively implemented. Each block could run at higher speed because only the fast local interconnections are used. Experimental results show an up to 28% performance improvement compared to the conventional FPGAs with small area overhead (around 2%).
Ranga Vemuri
FCCM2
2005 PAHLS: Towards Run-Time Synthesis for FPGAs
abstract
In this abstract, we have presented our research efforts toward the integration of physical synthesis with high level synthesis. By incorporating physical considerations into high level specification, we restrict computations and communications to geographic proximities while reserve the quality of the final result to a large extent within limited resources of FPGAs. We believe that the proposed methodology provides possible directions for synthesis unification of high level abstraction and lower level implementation, and is on the right track towards achieving a well-balanced (or even a globally optimum mapping, this is the long-run objective of PAHLS) synthesis result.
Renqiu Huang, Ranga Vemuri
FPL2
2005 A Novel Asynchronous FPGA Architecture Design and Its Performance Evaluation
abstract
This paper proposes GAPLA: a globally asynchronous locally synchronous programmable logic array architecture. The whole FPGA area is divided into locally synchronous blocks wrapped with asynchronous I/O interfaces. Data communications between synchronous blocks are controlled by 2-phase handshaking signals under bundled-data delay assumption. The size and shape of each locally synchronous block are programmable so that different modules in a design can be effectively implemented. By dividing the FPGA area into smaller blocks, the delays of long interconnect wires, which could easily dominate other delays in conventional FPGAs, only come into picture when there are communications between blocks. Therefore, each block could run at higher speed. The area overhead of adopting the GALS style in GAPLA architecture is estimated to be very small (about 7%). Experimental results show an up to 55% performance improvement compared to the conventional FPGAs.
Ranga Vemuri
FPL2
2005 Energy Management in Battery-Powered Sensor Networks with Reconfigurable Computing Nodes
abstract
In this work we have investigated the benefits of using reconfigurable computing (RC) nodes in sensor networks. We assumed that several sensor nodes are deployed randomly in a field, to form a sensor network and each sensor in the network sends its data in the form of packets to a single energy-rich sink node. We also assumed that each sensor node has reconfigurable fabric which can be configured by downloading a bitstream. In contrast to the contemporary work in energy management for sensor networks, we use an accurate analytical battery model to simulate the battery consumption of each node in the network. We have written several simulation models to study various sensor network parameters when the underlying nodes are adaptive in nature instead of traditional, non-adaptive processor based, fixed implementation. As the remaining battery-capacity of our RC based node decreases, it changes its behavior by reconfiguring itself to lower powered implementations successively, thereby extending the sensor network lifetime as a whole. Our results indicate that the network life is increased by up to five times and the number of packets generated by the sensor nodes and received at the sink node more than quadrupled for RC based nodes when compared to fixed processor based node implementation.
Jawad Khan, Ranga Vemuri
FPL2
2005 Accuracy driven performance macromodeling of feasible regions during synthesis of analog circuits
abstract
We propose an accuracy driven synthesis methodology for analog circuits. The proposed approach relies on macro-models for performance estimation and is thus orders of magnitude faster than simulation based synthesis techniques. Unlike existing macro-model based approaches, which use static models, our approach dynamically improves the accuracy of the model during synthesis to ensure true convergence. Our method is based on identifying and accurately modeling those regions in the design space where feasible designs lie. The identified feasible regions and their corresponding models are used in conjunction with an initially generated global model for performance estimation. Experimental results demonstrate that the proposed approach is able to synthesize designs using the enhanced performance models quickly, yet accurately.
Anuradha Agarwal, Glenn Wolfe, Ranga Vemuri
ACM Great Lakes Symposium on VLSI3
2005 Moment-driven coupling-aware routing methodology
abstract
An underdamped signal response with a number of overshoots and undershoots may lead to false switching and increased settling time delay. This 'ringing' effect adversely affects the signal quality at the output and becomes a source of major concern at multi-GHz frequencies, as the self and mutual inductance of interconnects start playing a crucial role in the performance of a circuit. Reduction in wire length or minimization of coupling capacitance, the stronghold in many earlier routing techniques, may produce a routing solution suitable only at sub-GHz frequencies. In this paper, we propose a routing methodology that accounts for inductive and capacitive parasitics (self and mutual) of the interconnects in its cost function through a combination of second and third order central moments. A trade-off between signal delay and amount of ringing, quantified by second and third order central moments respectively, has been made, which generates a routing solution with the best compromise between ringing and delay for each net under a monotone signal response.
Amitava Bhaduri, Ranga Vemuri
ACM Great Lakes Symposium on VLSI2
2005 LiPaR: A light-weight parallel router for FPGA-based networks-on-chip
abstract
Present day technology for ASICs supports Networks-on-Chip designs which can have 100 million gates on a single chip. The latest FPGAs can support only about 10 million gates to accomodate all logic and the associated routing. In order to implement a competitive NoC architecture in FP-GAs, the area occupied by the network should be kept to a minimum. This ensures that the maximum area can be utilized by the logic while maintaining the performance of the router network. Reducing area also reduces the power consumption. In this paper, we implement a parallel router which can support five simultaneous routing requests at the same time with an area overhead of only 352 Xilinx Virtex-II Pro FPGA slices (2. 57% of XC2VP30). We introduce optimizations in XY routing and decoding logic thereby gaining in area and performance. The header overhead is 8 bits per packet and the packet size can vary between 16 and 128 bits. We also implement a 3 x 3 mesh network with a total area overhead of 28% leaving 72% of the area available for the logic in a Virtex-II Pro XC2VP30 device. We characterize the router and several mesh networks for power and performance parameters.
Balasubramanian Sethuraman, Prasun Bhattacharya, Jawad Khan, Ranga Vemuri
ACM Great Lakes Symposium on VLSI4
2005 Hierarchical performance macromodels of feasible regions for synthesis of analog and RF circuits
abstract
Accurate performance modeling is essential for its usage in a circuit synthesis flow. Only a small fraction of the entire design space is occupied by designs with meaningful behavior and performance. In this work, we have focussed on modeling these feasible regions accurately in contrast with modeling the entire design space. Macromodels for the feasible regions were built hierarchically until the desired accuracy was achieved. An accuracy driven synthesis methodology is proposed to guide the identification of the feasible regions and dynamically enhance the performance of the macromodels. Dynamic performance modeling ensures true convergence of our synthesis approach as opposed to existing static macromodel based techniques. We applied the proposed methodology for modeling and synthesis of several analog and RF circuits and the results demonstrate that our approach yields highly accurate design solutions in a much smaller time compared to simulation based approaches.
Anuradha Agarwal, Ranga Vemuri
ICCAD2
2005 Layout-Aware RF Circuit Synthesis Driven by Worst Case Parasitic Corners
abstract
We propose a methodology for sizing radio-frequency circuits. Techniques for including layout information during circuit sizing have been presented. The aim of the proposed technique is to obtain parasitic closure at the post-layout validation stage. A two-step approach is adopted for achieving this goal. In the first step, the interconnect parasitic bounds are estimated. In the second step, the parasitic bounds are used to identify the worst case parasitics and the circuit is resized in presence of these parasitics. The proposed approach unlike existing layout-inclusive approaches achieves parasitic closure while not restricting the flexibility and quality of the physical layout. This methodology was applied on RF circuits like low noise amplifiers and the results demonstrate that our technique helps in obtaining robust parasitic aware design solutions.
Anuradha Agarwal, Ranga Vemuri
ICCD2
2004 Fast and accurate parasitic capacitance models for layout-aware
abstract
Considering layout effects early in the analog design process is becoming increasingly important. We propose techniques for estimating parasitic capacitances based on look-up tables and multi-variate linear interpolation. These models enable fast and accurate estimation of parasitic capacitances and are very suitable for use in a synthesis flow. A layout aware methodology for synthesis of analog CMOS circuits using these parasitic models is presented. Results indicate that the proposed synthesis system is fast as compared to a layout-inclusive synthesis approach.
Anuradha Agarwal, Hemanth Sampath, Veena Yelamanchili, Ranga Vemuri
DAC4
2004 An efficient algorithm for finding empty space for online FPGA placement
abstract
A fast and efficient algorithm for finding empty area is necessary for online placement, task relocation and defragmentation on a partially reconfigurable FPGA. We present an algorithm that finds empty area as a list of overlapping maximal rectangles. Using an innovative representation of the FPGA, we are able to predict possible locations of the maximal empty rectangles. Worst-case time complexity of our algorithm is O(xy) where x is the number of columns, y is the number of rows and x.y is the total number of cells on the FPGA. Experiments show that, in practice, our algorithm needs to scan less than 15% of the FPGA cells to make a list of all maximal empty rectangles.
Manish Handa, Ranga Vemuri
DAC2
2004 Accurate Estimation of Parasitic Capacitances in Analog Circuits
abstract
This paper presents efficient and accurate techniques for modeling parasitic capacitances in analog CMOS circuits. A layout aware synthesis flow using these parasitic models has been proposed. The fast parasitic estimation process replaces the time consuming steps of layout generation and extraction during synthesis. Results indicate that these models are extremely fast and accurate.
Anuradha Agarwal, Hemanth Sampath, Veena Yelamanchili, Ranga Vemuri
DATE4
2004 A Fast Algorithm for Finding Maximal Empty Rectangles for Dynamic FPGA Placement
abstract
In this paper, we present a fast algorithm for finding empty area on the FPGA surface with some rectangular tasks placed on it. We use a staircase data structure to report the empty area in the form of a list of maximal empty rectangles. We model the FPGA surface using an innovative encoding scheme that improves runtime and reduces memory requirement of our algorithm. Worst-case time complexity of our algorithm is O(xy) where x is number of columns, y is number of rows, and x.y is the total number of cells on the FPGA.
Manish Handa, Ranga Vemuri
DATE2
2004 Fast, Layout-Inclusive Analog Circuit Synthesis using Pre-Compiled Parasitic-Aware Symbolic Performance Models
abstract
We present a new methodology for fast analog circuit synthesis, based on the use of parameterized layout generators and symbolic performance models (SPMs) in the synthesis loop. Fast layout generation is achieved by using efficient parameterized procedural layout generators. Fast performance estimation is achieved by using pre-compiled SPMs, stored as efficient DDD-like structures called element coefficient diagrams. Techniques have been developed to include layout geometry effects in the SPMs. The accuracy and efficiency of the parasitic inclusion technique as well as the proposed methodology have been demonstrated by comparisons to traditional synthesis methods. The proposed methodology is used for the synthesis of opamps and filters and is demonstrated to achieve effective performance closure.
Mukesh Ranjan, Wim Verhaegen, Anuradha Agarwal, Hemanth Sampath, Ranga Vemuri, Georges Gielen
DATE5
2004 An Integrated Online Scheduling and Placement Methodology
Manish Handa, Ranga Vemuri
FPL2
2004 Analysis of a Hybrid Interconnect Architecture for Dynamically Reconfigurable FPGAs
Renqiu Huang, Manish Handa, Ranga Vemuri
FPL3
2004 A Dynamically Reconfigurable Asynchronous FPGA Architecture
Jayanthi Rajagopalan, Ranga Vemuri
FPL3
2004 An Efficient Battery-Aware Task Scheduling Methodology for Portable RC Platforms
Jawad Khan, Ranga Vemuri
FPL2
2004 A high level language for pre-layout extraction in parasite-aware analog circuit synthesis
abstract
This paper presents a high-level language MSL, for the specification of parameterized, topology-specific circuit extractors. Upon compilation, the MSL program yields an executable module which generates the extracted circuit containing parasitics, passive and active devices when given specific sizes. In contrast to traditional post-layout extraction, this is done without ever generating a layout. We call this pre-layout extraction. Pre-layout extraction is much faster than post-layout extraction and is highly suited for use in layout-aware circuit sizing programs. MSL can also be used for the specification of parameterized layout generators. Thus, although a concrete layout is never generated during pre-extraction, the extracted circuit is very much influenced by the symbolic placement and routing specified in the layout generation part of the MSL program. This ensures that the pre-layout extraction process yields the same results as post-layout extraction. Being a high-level language based approach, users can tune pre-layout extraction to a desired level of accuracy by modeling selected parasitics and ignoring others. This ability helps further speed up the circuit sizing process up to a factor varying from 2.5 to 4.5 compared to layout-inclusive synthesis methodologies.
Raoul F. Badaoui, Hemanth Sampath, Anuradha Agarwal, Ranga Vemuri
ACM Great Lakes Symposium on VLSI4
2004 Analysis and evaluation of a hybrid interconnect structure for FPGAs
abstract
In this paper, a cluster-based FPGA is proposed. The proposed FPGA has a hybrid interconnect structure which takes advantages of both mesh and tree topologies. We analyze the area and performance of proposed FPGA in terms of the needed switches by comparing with those of conventional FPGAs. We evaluate the proposed architecture on a series of benchmark designs. The experimental results show that the proposed model can significantly reduce the routing area, achieve high performance and admit more implementations of various designs at the price of a modest increase of switches required for that architecture.
Renqiu Huang, Ranga Vemuri
ICCAD2
2004 Adaptive sampling and modeling of analog circuit performance parameters with pseudo-cubic splines
abstract
Many approaches to analog performance parameter macro modeling have been investigated by the research community. These models are typically derived from discrete data obtained from circuit simulation using numerous input combinations of component sizes for a given circuit topology. The simulations are computationally intensive, therefore it is advantageous to reduce the number of simulations necessary to build an accurate macro model. We present a new algorithm for adaptively sampling multi-dimensional black box functions based on Duchon pseudo-cubic splines. The splines readily and accurately model high dimensional functions based on discrete unstructured data and require no tuning of parameters as seen in many other interpolation methods. The adaptive sampler, in conjunction with pseudo-cubic splines, is used to accurately model various analog performance parameters for an operational amplifier topology using fewer sample points than traditional gridded and quasi-random sampling methodologies.
Ranga Vemuri, Glenn Wolfe
ICCAD1
2004 Simultaneous Scheduling, Binding and Layer Assignment for Synthesis of Vertically Integrated 3D Systems
abstract
Three-dimensional vertically integrated systems allow active devices to be placed on multiple device layers. In recent years, a number of research efforts have addressed physical synthesis issues for such systems. Such efforts showed a significant reduction in interconnect lengths. In order to effectively synthesize designs for 3D systems, it is necessary to take layer assignment for resources into consideration at higher levels of the design abstraction. We address the layer assignment problem as a part of a physical aware behavioral synthesis flow. We propose a 0-1 linear program formulation to perform simultaneous and optimal scheduling, binding and layer assignment for synthesizing designs for three-dimensional vertically integrated systems. The objective is to minimize inter-stratal via and the interconnect length in the critical path while taking thermal gradient between layers into account (which has been shown to be of particular concern for 3D systems). Floorplanning is performed for the synthesized design in order to estimate interconnect lengths. Results show a reduction of approximately 37% in total interconnect lengths on an average, compared to a traditional two-dimensional implementation when 2-5 layer implementations are examined.
Madhubanti Mukherjee, Ranga Vemuri
ICCD2
2004 Hardware Assisted Two Dimensional Ultra Fast Placement
abstract
Summary form only given. Placement time is an overhead on the application execution time in an online placement system. In a partially reconfigurable system, the inherent parallelism of the reconfigurable hardware can be explored to speed up the placement process. We present three different architectures for two dimensional online placement. Each architecture makes different trade-offs between area usage, memory requirement and execution time. These architectures are capable of achieving very fast placement while using a very small number of hardware resources.
Manish Handa, Ranga Vemuri
IPDPS2
2004 Forward-Looking Macro Generation and Relational Placement During High Level Synthesis to FPGAs
abstract
Summary form only given. Incorporating physical information into earlier architectural and logic synthesis stages is highly desirable since it allows more realistic exploration of the design space and the generation of solutions with predictable metrics. We present a forward-looking synthesis methodology in which we weigh all nets in the control data flow graph (CDFG) according to their criticality. We cluster operations in the CDFG into macros while satisfying logical and physical constraints. We perform relational placement on these macros. We have evaluated the proposed approach using a set of benchmark designs by comparing it with the results of a traditional synthesis flow. The results show that our methodology achieves up to 26% improvement in clock frequency without any area overhead, and average 12.7% improvement in critical path delay with no or little place-and-route time overhead.
Renqiu Huang, Ranga Vemuri
IPDPS2
2004 A two-layer library-based approach to synthesis of analog systems from VHDL-AMS specifications
abstract
This paper presents a synthesis methodology for analog systems described using VHDL-AMS language. Synthesis produces net-lists of analog components that are selected from a library, and sized so that specified objectives (like AC response, signal to noise ratio, dynamic range, area) are optimized. The gap between abstract specifications and implementations is bridged using a two-layered methodology. The first layer is architecture generation. The second layer is component synthesis and constraint transformation. Architecture generation employs the branch-and-bound algorithm to create architectural alternatives for a system. Component synthesis and constraint transformation use a directed interval based genetic algorithm that operates on parameter ranges. The performance estimation engine embeds technology process parameters, SPICE models for basic circuits, and symbolic composition equations for basic structural configurations. The paper discusses the VHDL-AMS subset for synthesis. The subset offers the composition semantics. As a result, specifications offer sufficient insight into the system structure to allow automated architecture generation. To justify the flexibility of the methodology, the paper presents results for three case studies, a signal conditioning system, a filter, and an analog to digital converter. Experiments show that constraint-satisfying designs can be synthesized in a short time, at a low cost, and without requesting broad knowledge on analog circuits.
Alex Doboli, Nagu R. Dhanwada, Adrián Núñez-Aldana, Ranga Vemuri
ACM Trans. Design Autom. Electr. Syst.4
2003 A Novel Synthesis Strategy Driven by Partial Evaluation Based Circuit Reduction for Application Specific DSP Circuits
abstract
Traditional application specific synthesis systems for DSP rely on a predesigned library of components. Designs in the DSP domain often involve constant operands. Using off-the-shelf library components for such designs can be wasteful in terms of area, power and timing. Operations involving constant operands are possible candidates for reduction based on partial evaluation. Classical logic synthesis is often incapable of performing reductions because of component sharing (typically decided during high level synthesis) between regular operations and those involving constant operands, which prevent minimizations during this phase. We propose a methodology for performing on-demand component reduction using partial evaluation during synthesis of application specific DSP circuits. The simplified components have better characteristics compared to their unreduced counterparts in terms of both delay and power. Use of reduced components in the synthesis loop showed a system wide improvement in performance, area and power for the synthesized designs for a variety of benchmark DSP circuits.
Madhubanti Mukherjee, Ranga Vemuri
ICCD2
2003 MSL: A High-Level Language for Parameterized Analog and Mixed Signal Layout Generators
Hemanth Sampath, Ranga Vemuri
VLSI-SOC2
2003 Adaptive Sampling and Modeling of Analog Circuit Performance Parameters
Glenn Wolfe, Mengmeng Ding, Ranga Vemuri
VLSI-SOC3
2003 Behavioral modeling for high-level synthesis of analog and mixed-signal systems from VHDL-AMS
abstract
High-level synthesis is highly demanded for managing the complexity of analog and mixed-signal system designs. However, synthesis methods are currently in their infancy. The absence of a high-level specification notation is an important limitation for the development of efficient synthesis methods. This paper presents a behavioral model and a VHDL-AMS subset for high-level synthesis of analog and mixed-signal systems. The model (named aBlox) offers a composition semantics for functionality description and an orthogonal declarative mechanism for expressing the performance requirements of a system. The model was developed after analyzing a large number of systems for telecommunication, signal processing, control engineering, and analog computing. The model expresses the meaning of: 1) analog and digital data; 2) continuous and event-driven functionality (behavior); 3) analog performance attributes; and 4) analog-digital interactions. The aBlox model serves as a foundation for defining a semantically sound VHDL-AMS subset for synthesis. Also, the VHDL-AMS subset is identified so that its constructs can be mapped to architectures of circuits. We introduce several restrictions to the VHDL-AMS instructions, such that their semantics match that of the aBlox model. To motivate the usefulness of the model and the VHDL-AMS subset, we present a case study that uses VHDL-AMS inputs.
Alex Doboli, Ranga Vemuri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Exploration-based high-level synthesis of linear analog systems operating at low/medium frequencies
abstract
This paper presents a methodology for high-level synthesis of continuous-time linear analog systems. Synthesis results are architectures of op-amps, sized resistors and capacitors such that their ac behavior and total silicon area are optimized. Bounds for op-amp dc gain, unity-gain frequency, input, and output impedances are found as a byproduct of synthesis. Subsequently, a circuit synthesis tool can be used to synthesize the op-amps of an architecture. The paper details the architecture generation technique. Architecture generation produces alternative architectures for a system specification using the tabu search heuristic. Its main advantages over traditional methods is that it is application independent, does not require a library of block connection patterns, and is simple to implement. The paper also discusses the hierarchical, two-step parameter optimization that guides architecture generation. Experiments showed that linear analog systems operating at low/medium frequencies (like telecommunication systems and filters) can be synthesized in a reasonably long time and with reduced effort.
Alex Doboli, Ranga Vemuri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Extraction and use of neural network models in automated synthesis of operational amplifiers
abstract
Fast and accurate performance estimation methods are essential to automated synthesis of analog circuits. Development of analog performance models is difficult due to the highly nonlinear nature of various analog performance parameters. This paper presents a neural network-based methodology for creating fast and efficient models for estimating the performance parameters of CMOS operational amplifier topologies. Effective methods for generation and use of the training data are proposed to enhance the accuracy of the neural models. The efficiency and accuracy of the resulting performance models are demonstrated via their use in a genetic algorithm-based circuit synthesis system. The genetic synthesis tool optimizes a fitness function based on user-specified performance constraints. The performance parameters of the synthesized circuits are validated by SPICE simulations and compared with those predicted by the neural network models. Experimental studies demonstrate that neural network modeling is an effective, fast, and accurate methodology for performance estimation.
Glenn Wolfe, Ranga Vemuri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 A Functional Specification Notation for Co-Design of Mixed Analog-Digital Systems
abstract
This paper discusses aBlox - a specification notation for high-level synthesis of mixed-signal systems. aBlox addresses three important aspects of mixed-signal system specification: (1) description of functionality and (2) performance issues and (3) expression of analog-digital interactions. The semantics of aBlox embeds concepts and rules of a functional computational model, and uses a declarative style to denote performance elements. The paper shows some mixed-signal specifications that we developed in aBlox. Finally, we describe a high-level analog synthesis experiment that used aBlox specifications as inputs.
Alex Doboli, Ranga Vemuri
DATE2
2002 iPACE-V1: A Portable Adaptive Computing Engine for Real Time Applications
Jawad Khan, Manish Handa, Ranga Vemuri
FPL3
2002 An efficient register optimization algorithm for high-level synthesis from hierarchical behavioral specifications
abstract
We address the problem of register optimization that arises during high-level synthesis from modular hierarchical behavioral specifications. Register optimization is the process of grouping carriers such that each group can be safely allocated to a hardware register. Global register optimization by inline expansion involves flattening the module hierarchy and using a heuristic register optimization procedure on the flattened description. Although inline expansion yields a near-optimal number of registers, it is very time consuming due to the large number of carrier compatibility relationships that must be considered. We present an efficient register optimization algorithm that achieves nearly the same effect of inline expansion without actually inline expanding. The distinguishing feature of the proposed algorithm is that it employs a hierarchical optimization phase which effectively exploits the properties of the module call graph and information gathered during local carrier lifecycle analysis of each module. Experimental results on a number of benchmarks show that the proposed algorithm produces nearly the same number of registers as inline expansion based global optimization and is faster by a factor of 7.0.
Ranga Vemuri, Srinivas Katkoori, Meenakshi Kaul, Jay Roy
ACM Trans. Design Autom. Electr. Syst.1
2002 Hardware-software partitioning and pipelined scheduling of transformative applications
abstract
Transformative applications are computation intensive applications characterized by iterative dataflow behavior. Typical examples are image processing applications like JPEG, MPEG, etc. The performance of embedded hardware-software systems that implement transformative applications can be maximized by obtaining a pipelined design. We present a tool for hardware-software partitioning and pipelined scheduling of transformative applications. The tool uses iterative partitioning and pipelined scheduling to obtain optimal partitions that satisfy the timing and area constraints. The partitioner uses a branch and bound approach with a unique objective function that minimizes the initiation interval of the final design. We present techniques for generation of good initial solution and search-space limitation for the branch and bound algorithm. A candidate partition is evaluated by generating its pipelined schedule. The scheduler uses a novel retiming heuristic that optimizes the initiation interval, number of pipeline stages, and memory requirements of the particular design alternative. We evaluate the performance of the retiming heuristic by comparing it with an existing technique. The effectiveness of the entire tool is demonstrated by a case study of the JPEG image compression algorithm. We also evaluate the run time and design quality of the tool by experimentation with synthetic graphs.
Karam S. Chatha, Ranga Vemuri
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Integrated High-Level Synthesis and Power-Net Routing for Digital Design under Switching Noise Constraints
abstract
This paper presents a CAD methodology and a tool for high-level synthesis (HLS) of digital hardware for mixed analog-digital chips. In contrast to HLS for digital applications, HLS for mixed-signal systems is mainly challenged by constraints, such as digital switching noise (DSN), that are due to the analog circuits. This paper discusses an integrated approach to HLS and power net routing for effectively reducing DSN. Motivation for this research is that HLS has a high impact on DSN reduction, however, DSN evaluation is very difficult at a high level. Integrated approach also employs an original method for fast evaluation of DSN and an algorithm for power net routing and sizing. Experiments showed that our combined binding and scheduling method produces better results than traditional HLS techniques. Finally, DSN evaluation using the proposed algorithm can be significantly faster than SPICE simulation.
Alex Doboli, Ranga Vemuri
DAC2
2001 Behavioral Partitioning in the Synthesis of Mixed Analog-Digital Systems
abstract
Synthesis of mixed-signal designs from behavioral specifications must address analog-digital partitioning. In this paper, we investigate the issues in mixed-signal behavioral partitioning and design space exploration for signal-processing systems. We begin with the system behavior specified in an intermediate format called the Mixed Signal Flow Graph, based on the time-amplitude characterization of signals. We present techniques for analog-digital behavioral partitioning of the MSFG, and performance estimation of the technology-mapped analog and digital circuits. The partitioned solution must satisfy constrants on imposed by the target field programmable mixed-signal architecture on avaialable configurable resources, available data converters, their resolution and speed, and IO pins. The quality of the solution is evaluated based on two metrics, namely feasibility and performance. The former is a measure of the validity of the solution with respect to the architectural constraints. The latter measures the performance of the system based on bandwidth/speed and noise.
Sree Ganesan, Ranga Vemuri
DAC2
2001 A regularity-based hierarchical symbolic analysis method for large-scale analog networks
abstract
Summary form only given. The main challenge for any symbolic analysis method is the exponential size of the produced symbolic expressions (10/sup 11/ terms for an op amp). Current research considers two ways of handling this limitation: approximation of symbolic expressions and hierarchical methods. Approximation methods retain only the significant terms of the symbolic expressions and eliminate the insignificant ones. The difficulty, however, lies in identifying what terms to eliminate and what the resulting approximation error could be. Hierarchical methods tackle the symbolic analysis problem in a divide-and-conquer manner. They consider only one part of the global network at a time and then recombine partial expressions for finding overall symbolic formulas. Existing hierarchical methods have a main limitation in that they are not feasible for addressing networks that are built of tightly coupled blocks, i.e. operational amplifiers. The originality of our research stems from exploiting regularity aspects for addressing the exponential size of produced symbolic expressions. As a result, polynomial-size models are obtained for any network, including networks formed of tightly coupled blocks. Two kinds of regularity aspects were identified.
Alex Doboli, Ranga Vemuri
DATE2
2001 Hierarchical memory mapping during synthesis in FPGA-based reconfigurable computers
abstract
One step in the synthesis for FPGA-based Reconfigurable Computers (RCs) involves mapping the design data structures onto the physical memory banks available in the hardware. The advent of Xilinx Virtex-style FPGAs and of hierarchical memory schemes on reconfigurable boards introduced an added complexity to this mapping. The new RC boards offer a wealth of memory banks many of them on-chip (such as the BlockRAMs available in the Virtex architecture) and many of them offering variable number of ports and several depth/width configurations. Along with the external RAMs, a hierarchy of memories with varying access performances are available in a reconfigurable computer. It becomes critical to perform a good mapping to achieve optimal design performance. This paper presents an automatic memory mapping methodology which takes into account: the number of words and word size of design data segments and physical memory banks, number of ports on the banks, access latency of the banks, proximity of the banks to the processing unit, life cycle analysis of data segments, and it also incorporates configuration selection from the multiple configurations available in BlockRAMs of Virtex series FPGAs. In the case of multiple processing elements on board, the paper also provides a framework in which the task of memory mapping interacts with spatial partitioning to provide the best implementation.
Iyad Ouaiss, Ranga Vemuri
DATE2
2001 On the verification of synthesized designs using automatically generated transformational witnesses
abstract
Summary form only given. The authors present a new methodology for verifying the synthesized designs, and for debugging the software implementation of high-level synthesis algorithms. The methodology is based on a a set of 7 RTL transformations which are able to emulate the effect of many scheduling and resource allocation algorithms.
Elena Teica, Rajesh Radhakrishnan, Ranga Vemuri
DATE3
2001 Memory Synthesis for FPGA-Based Reconfigurable Computers
Amit Kasat, Iyad Ouaiss, Ranga Vemuri
FPL3
2001 Global memory mapping for FPGA-based reconfigurable systems
abstract
Synthesizing designs for FPGA-based reconfigurable systems involves the task of mapping variables and data structures of the application onto RAMs of the reconfigurable board. The variety in types and performance of onboard and on-chip RAMs, their proximity to the processing units, and the interconnection scheme of the reconfigurable system, all contribute to an intricate memory mapping problem. An intelligent memory assignment minimizes the total latency of the design and the interconnection requirements due to memory accesses. A complete Integer Linear Programming (ILP) formulation of the problem results in an optimized memory mapping; however, the formulation is complex and takes a very long time to produce a solution. In order to efficiently solve the problem, the concept of global/detailed memory mapping is introduced in this paper. An ILP formulation of the global mapping process is described. This formulation is simpler and faster than the complete formulation, and it leaves the task of detailed mapping to a post-ILP tool that does not affect the optimality of the memory assignment. As a result, larger designs can be handled at a faster rate and more constraints can be introduced to the formulation.
Iyad Ouaiss, Ranga Vemuri
IPDPS2
2001 Continous Wavelet Transform on Reconfigurable Meshes
abstract
Wavelet transforms have proven to be useful tools for several applications, including signal analysis, signal coding, and image compression. In this paper, faster parallel algorithms for computing the continuous wavelet transform are designed for reconfigurable meshes. An -time algorithm for computing the continuous wavelet transform with signals and an integer grid on a 3-D reconfigurable mesh is proposed, where is the number of bits used to represent the values in calculation. A constant-time algorithm 3-D reconfigurable mesh is also proposed. To the best knowledge of the author, this is the first constanttime algorithm for continuous wavelet transform on any parallel architecture.
Yi Pan 0001, Jie Li 0002, Ranga Vemuri
IPDPS3
2001 Theorem Proving Guided Development of Formal Assertions in a Resource-Constrained Scheduler for High-Level Synthesis
Naren Narasimhan, Elena Teica, Rajesh Radhakrishnan, Sriram Govindarajan, Ranga Vemuri
Formal Methods Syst. Des.5
2001 Fine-grained and coarse-grained behavioral partitioning with effective utilization of memory and design space exploration for multi-FPGA architectures
abstract
Reconfigurable computers (RCs) host multiple field programmable gate arrays (FPGAs) and one or more physical memories that communicate through an interconnection fabric. State-of-the-art RCs provide abundant hardware and storage resources, but have tight constraints on FPGA pin-out and inter-FPGA interconnection resources. These stringent constraints are the primary impediment for multi-FPGA partitioning tools to generate high-quality designs, in this paper, we present two integrated partitioning and synthesis approaches for RCs. The first approach involves fine-grained partitioning of a scheduled data-flow graph (DFG, or an operation graph), and the second involves a coarse-grained partitioning of an unscheduled control data flow graph (CDFG, or a block graph). A hardware design space exploration engine is integrated with the block graph partitioner that dynamically contemplates multiple schedules during partitioning. The novel feature in the partitioning approaches is that the physical memory in the RC is effectively used to alleviate the FPGA pin-out and inter-FPGA interconnection bottle-neck. Several experiments have been conducted, targeting commercial multi-FPGA boards, to compare the two partitioning approaches, and detailed summaries are presented.
Sriram Govindarajan, Ranga Vemuri
IEEE Trans. Very Large Scale Integr. Syst.3
2001 Guest editorial reconfigurable and adaptive VLSI systems
Ranga Vemuri, Rajesh K. Gupta 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2000 Technology Mapping and Retargeting for Field-Programmable Analog Arrays
abstract
Rapid prototyping followed by technology retargeting provides a fast and cost-effective approach to analog system synthesis. Field-programmable analog arrays (FPAAs) enable rapid implementation of a function-compliant prototype, while technology retargeting converts the functional FPAA prototype to an ASIC. We first address the FPAA technology mapping problem. A novel structural approach based on hierarchical pattern matching and covering is employed to map the analog behavior onto the FPAA. We then address the issues of technology retargeting and design reuse, and present out FPAA-ASIC retargering strategy. We present experiments and a design example for FPAA technology mapping and retargeting.
Sree Ganesan, Ranga Vemuri
DATE2
2000 An Integrated Temporal Partitioning and Partial Reconfiguration Technique for Design Latency Improvement
abstract
Partially reconfigurable processors provide the unique ability by which a part of the device can be reconfigured, while the remaining part is still operational. In this paper, we present a novel partitioning methodology that temporally partitions a design for such a partially reconfigurable processor and improves design latency by minimizing reconfiguration overhead. This is achieved by overlapping execution of one temporal partition with the reconfiguration of another, using the processors partial reconfiguration capability. We have incorporated block-processing in the partitioning framework for reducing reconfiguration overhead of partitioned designs. A highlight of our partitioner is it's ability to handle loops and conditional constructs in the input specification. The proposed methodology was tested on several examples on the Xilinx 6200 FPGA. The results show significant reduction in the design latency, leading to a considerable speed-up due to partial reconfiguration.
Satish Ganesan, Ranga Vemuri
DATE2
2000 Improving the Schedule Quality of Static-List Time-Constrained Scheduling
abstract
Summary form only given. The most compelling reason for High-Level Synthesis (HLS) to be accepted in the state-of-the-art CAD flow is its ability to perform design space exploration. Design space exploration requires efficient scheduling techniques that have a low complexity and yet produce good quality schedules. The Time-Constrained Scheduling (TCS) problem minimizes the number of functional units required to schedule a particular Data Flow Graph (DFG) within a specified number of time steps. Over the past few years a number of techniques have been proposed to solve the TCS problem. Heuristic list scheduling algorithms have been widely used for their low-complexity and good performance. The complexity of a dynamic-list scheduling algorithm, such as the Force Directed Scheduling (FDS), is /spl Theta/(T*N/sup 2/), where T is the time constraint and N is the number of operations. Static-list scheduling algorithms are the least complex among the known class of scheduling techniques with a linear time complexity of /spl Theta/(T*N). Typically, static-list scheduling algorithms, in order to maintain low-complexity, do not perform any look-ahead like that of FDS. The drawback is that, static-list scheduling algorithms may not generate high-quality schedules. However, the proposed static-list algorithm presented here incorporates a novel topological clustering technique which acts as the look-ahead mechanism without any computational overhead.
Sriram Govindarajan, Ranga Vemuri
DATE2
2000 Efficient Resource Arbitration in Reconfigurable Computing Environments
abstract
In a multi-FPGA synthesis system, ideally the designer has only an abstract view of the board architecture. This abstract modeling of the underlying reconfigurable computer poses complex challenges to the synthesis and partitioning tools. Since the design specification is not constrained by the number of memory segments on the board or the number of pins between FPGAs, it is difficult for the CAD tools to transform the design into one that maps onto the multi-FPGA board. This paper describes an arbitration mechanism that bridges the abstraction between the implicit design and the reconfigurable architecture. Since this mechanism allows such architecture abstraction between the design and the board, it becomes easier to port a design from one target architecture to another. This arbitration mechanism introduces very little overhead in terms of area and delay. It has been used in data-dominated applications; in this paper fast Fourier transform (FFT) is shown as an illustrative example.
Iyad Ouaiss, Ranga Vemuri
DATE2
2000 A heuristic technique for system-level architecture generation from signal-flow graph representations of analog systems
abstract
This paper presents a heuristic technique for automatically generating different architectures for an analog system. The AG iteratively produces various system net-lists as distinct implementations can realize the signal processing and flow in a system. Area and power for resulting net-lists are rapidly evaluated with High-Level Performance Estimator (HPE), a simplified estimation module. The AG algorithm is simple to implement. It does not require an extensive pattern library as traditional AG techniques do.
Alex Doboli, Nagu R. Dhanwada, Ranga Vemuri
ISCAS3
2000 Scheduling for low power under resource and latency constraints
abstract
We extend the Force-Directed List Scheduling (FDLS) algorithm proposed by Paulin and Knight (1989). In any time step, if the number of ready operations exceeds the available functional resources, then some operations must be deferred. The concept of "force" introduced by Paulin and Knight captures the effect of deferring an operation on the schedule length: larger the force, lower the likelihood of schedule length increase due to the operation's deferral. We develop a power cost function that captures the effect of an operation's deferral on the total power consumption of the design. The novelty of the work lies in heuristically determining the "best" time-step for an operation such that the overall power consumption is minimized without sacrificing the design throughput. The power-delay cost function proposed at the operation-level facilitates such an exploration. Experimental results show power savings of up to 60% with an average power savings of 23% at datapath level, 3% at the controller level, and 14% at the design-level.
Srinivas Katkoori, Ranga Vemuri
ISCAS2
2000 Automated Correctness Condition Generation for Formal Verification of Synthesized RTL Designs
Nazanin Mansouri, Ranga Vemuri
Formal Methods Syst. Des.2
1999 Automatic Constraint Transformation with Integrated Parameter Space Exploration in Analog System Synthesis
abstract
In this paper, we present a constraint transformation and topology selection methodology that explores the system level parameter space to compute acceptable regions in the component parameter space. The search process of an underlying circuit synthesis tool could be confined to these regions of valid solutions. Experimental results showing the impact of parameter space exploration at a higher level on analog circuit synthesis are presented demonstrating the effectiveness of this technique.
Nagu R. Dhanwada, Adrián Núñez-Aldana, Ranga Vemuri
ASP-DAC3
1999 Behavioral Synthesis of Analog Systems Using Two-layered Design Space Exploration
abstract
This paper presents a novel approach for synthesis of analog systems from behavioral VHDL-AMS specifications.We implemented this approach in the VASE behavioral-synthesis tool.The synthesis process produces a netlist of electronic components that are selected from a component library and sized such that the overall area is minimized and the rest of the performance constraints such as power, slew-rate, bandwidth, etc. are met.The gap between system level specifications and implementations is bridged using a hierarchically-organized, design-space exploration methodology.Our methodology performs a two-layered synthesis, the first being architecture generation, and the other component synthesis and constraint transformation.For architecture generation we suggest a branch-and-bound algorithm, while component synthesis and constraint transformation use a Genetic Algorithm based heuristic method.Crucial to the success of our exploration methodology is a fast and accurate performance estimation engine that embeds technology process parameters, SPICE models for basic circuits and performance composition equations.We present a telecommunication application as an example to illustrate our synthesis methodology, and show that constraint-satisfying designs can be synthesized in a short time and with a reduced designer effort.
Alex Doboli, Adrián Núñez-Aldana, Nagu R. Dhanwada, Sree Ganesan, Ranga Vemuri
DAC5
1999 An Automated Temporal Partitioning and Loop Fission Approach for FPGA Based Reconfigurable Synthesis of DSP Applications
abstract
We present an automated temporal partitioning and loop transformation approach for developing dynamically reconfigurable designs starting from behavior level specifications.An Integer Linear Programming (ILP) model is formulated to achieve near-optimal latency designs.We, also present a loop restructuring method to achieve maximum throughput for a class of DSP applications.This restructuring transformation is performed on the temporally partitioned behavior and results in near-optimization of throughput.We discuss eficient memory mapping and address generation techniques for the synthesis of reconfigurable designs.A Case study on the Joint Photographic Experts Group (JPEG) image compression algorithm demonstrates the egectiveness of our approach.
Meenakshi Kaul, Ranga Vemuri, Sriram Govindarajan, Iyad Ouaiss
DAC2
1999 Hierarchical Constraint Transformation Using Directed Interval Search for Analog System Synthesis
abstract
In this paper, we present a hierarchical approach for constraint transformation. The important features of this are: a genetic algorithm (GA) based search engine that computes design parameter ranges, a hierarchically organized characterization mechanism based on the concept of directed intervals that assists the search engine and an analog performance estimator. Experiments were conducted comparing the hierarchical approach with a flat bottom-up one. The results obtained demonstrate the effectiveness of the former approach. Experimental results highlighting the impact of using the characterization information within the constraint transformation process are also presented.
Nagu R. Dhanwada, Adrián Núñez-Aldana, Ranga Vemuri
DATE3
1999 A VHDL-AMS Compiler and Architecture Generator for Behavioral Synthesis of Analog Systems
abstract
This paper presents a complete method for automatically translating VHDL-AMS behavioral-specifications of analog systems into op amp level net-lists of library components. We discuss the three fundamental aspects, that pertain to any behavioral synthesis environment the specification language, the rules for compiling language constructs into a technology-independent, intermediate representation, and the synthesis (mapping) of representations to net-lists (topologies) of library components, so that performance constraints are satisfied. We motivate the effectiveness of the method by presenting our synthesis results for 5 examples.
Alex Doboli, Ranga Vemuri
DATE2
1999 Temporal Partitioning combined with Design Space Exploration for Latency Minimization of Run-Time Reconfigured Designs
abstract
We present combined temporal partitioning and design space exploration techniques for synthesizing behavioral specifications for run-time reconfigurable processors. Design space exploration involves selecting a design point for each task from a set of design points for that task to achieve latency minimization of partitioned solutions. We present an iterative search procedure that uses a core ILP (integer linear programming) technique, to obtain constraint satisfying solutions. The search procedure explores different regions of the design space while accomplishing combined partitioning and design space exploration. A case study of the DCT (discrete cosine transform) demonstrates the effectiveness of our approach.
Meenakshi Kaul, Ranga Vemuri
DATE2
1999 Accounting for Various Register Allocation Schemes During Post-Synthesis Verification of RTL Designs
abstract
This paper reports a formal methodology for verifying a broad class of synthesized register-transfer-level (RTL) designs by accommodating various register allocation/optimization schemes commonly found in high-level synthesis tools. Performing register optimization as part of synthesis process implies that the mapping between the specification variables and RTL registers is not bijective. We propose a formalization of dynamic variable-register mapping, and techniques based on symbolic analysis and higher-order logic theorem proving for verifying synthesized RTL designs. The proposed verification methodology has been successfully implemented using the PVS theorem prover.
Nazanin Mansouri, Ranga Vemuri
DATE2
1999 An Analog Performance Estimator for Improving the Effectiveness of CMOS Analog Systems Circuit Synthesis
abstract
Critical to the automation of analog circuit systems is the estimation process of performance parameters which are used to guide the topology selection and circuit sizing processes. This paper presents a methodology to improve the effectiveness of the CMOS analog system circuit synthesis search process by developing an Analog Performance Estimator (APE) tool. APE is capable of accepting the design parameters of an analog circuit and determine its performance parameters along with anticipated sizes of all the circuit elements. The APE is structured as a hierarchical estimation engine containing performance models of analog circuits at various levels of abstraction.
Adrián Núñez-Aldana, Ranga Vemuri
DATE2
1999 Task-Level Partitioning and RTL Design Space Exploration for Multi-FPGA Architectures
abstract
This paper presents SPADE, a system for partitioning designs onto multi-FPGA architectures. The input to SPADE is a task graph, that is composed of computational tasks, memory tasks and the communication and synchronization between tasks. SPADE consist of an iterative partitioning engine, an architectural constraint evaluator, and a throughput optimization and RTL design space exploration heuristic. We show how various architectural constraints can be effectively handled using an iterative partitioning engine.
Vinoo Srinivasan, Ranga Vemuri
FCCM2
1999 Throughput Optimization with Design Space Exploration During Partitioning for Multi-FPGA Architectures
abstract
No abstract available.
Vinoo Srinivasan, Ranga Vemuri
FPGA2
1999 Hierarchical Scheduling in High Level Synthesis Using Resource Sharing Across Nested Loops
abstract
This paper presents a resource-constrained scheduling algorithm for hierarchical behavioral specifications containing nested loops. The algorithm attempts to share resources across levels, to schedule operations that belong to different levels of the nested loop structures in the specifications as well as operations that belong to the same level. We compare the results of scheduling using our algorithm with those obtained using traditional list scheduling with no sharing of resources among different levels of the specification. These results show an average improvement of 23.47% in terms of number of control steps.
Abhijit Ghosh, Sandeep K. Lodha, Ranga Vemuri
Great Lakes Symposium on VLSI3
1999 Accurate Resource Estimation Algorithms for Behavioral Synthesis
abstract
Given a scheduled data flow graph the functional, storage, and interconnect (multiplexors) resources are analytically estimated taking into account the effects of post-scheduling tasks. Complexity of the controller implementation is also estimated. The novelty of this work lies in predicting the effects of the post-scheduling task on the final amount of resources, the effects of data path resource optimization on the controller complexity. Experimental results show high correlation between estimated and actual numbers.
Srinivas Katkoori, Ranga Vemuri
Great Lakes Symposium on VLSI2
1999 A Methodology for Rapid Prototyping of Analog Systems
abstract
We present a methodology for rapid prototyping of linear time-invariant analog systems. The prototyping hardware is composed of field-programmable analog arrays (FPAAs) to enable rapid evaluation and validation of analog designs. Starting with a signal flow graph description of the system, a library-based technology mapping phase produces FPAA designs optimized for area. The technology mapper then explores the design space by performing gain distribution. Technology mapping is followed by the placement and routing phase that generated the physical layout on the target single-segment array-based FPAA architecture. We employ an integrated place and route approach in order to guarantee routability and performance.
Sree Ganesan, Ranga Vemuri
ICCD2
1999 Formal Verification of Synthesized Analog Designs
abstract
We present an approach for formal verification of the DC and low frequency behavior of synthesized analog designs containing linear components and components whose behavior can be represented by piecewise linear models. A formal model of the structural description of a synthesized design is extracted from the sized component netlist produced by the synthesis tool, in terms of characteristic behavior of the components and various voltage and current laws. For the implementation to be correct, it must imply a formal model extracted from a user given behavior specification. Circuit implementation and expected behavior are both modeled in the PVS higher-order logic proof checker as linear functions and the PVS decision procedures are used to prove the implication.
Abhijit Ghosh, Ranga Vemuri
ICCD2
1998 Optimal Temporal Partitioning and Synthesis for Reconfigurable Architectures
abstract
We develop a 0-1 non-linear programming (NLP) model for combined temporal partitioning and high-level synthesis from behavioral specifications destined to be implemented on reconfigurable processors. We present tight linearizations of the NLP model. We present effective variable selection heuristics for a branch and bound solution of the derived linear programming model. We show how tight linearizations combined with good variable selection techniques during branch and bound yield optimal results in relatively short execution times.
Meenakshi Kaul, Ranga Vemuri
DATE2
1998 Hardware Software Partitioning with Integrated Hardware Design Space Exploration
abstract
This paper presents an integrated approach to hardware software partitioning and hardware design space exploration. We propose a genetic algorithm which performs hardware software partitioning on a task graph while simultaneously contemplating various design alternatives for tasks mapped to hardware. We primarily deal with data dominated designs typically found in digital signal processing and image processing applications. A detailed description of various genetic operators is presented. We provide results to illustrate the effectiveness of our integrated methodology.
Vinoo Srinivasan, Shankar Radhakrishnan, Ranga Vemuri
DATE3
1998 An Effective Design System for Dynamically Reconfigurable Architectures
abstract
The SPARCS system is an integrated partitioning and synthesis environment for reconfigurable architectures. In this paper, we use the Joint Photographic Experts Group (JPEG) image compression algorithm as a design example to demonstrate the effectiveness of dynamic reconfiguration achieved using SPARCS. We present a typical design process using the SPARCS system consisting of temporal partitioning, spatial partitioning, and design synthesis. The results, obtained on a commercial RC architecture, show that the multiply-reconfigured version of the JPEG compression algorithm achieves reasonable improvement in execution times compared to the one-time configured version.
Sriram Govindarajan, Iyad Ouaiss, Meenakshi Kaul, Vinoo Srinivasan, Ranga Vemuri
FCCM5
1998 A Methodology for Automated Verification of Synthesized RTL Designs and Its Integration with a High-Level Synthesis Tool
Nazanin Mansouri, Ranga Vemuri
FMCAD2
1998 Theorem proving guided development of formal assertions in a resource-constrained scheduler for high-level synthesis
abstract
A formal specification and a proof of correctness of the widely-used force-directed list scheduling (FDLS) algorithm for resource-constrained scheduling in high-level synthesis systems is presented. The proof effort is conducted using a higher-order logic theorem prover. During the proof effort many interesting properties of the FDLS algorithm are discovered. These properties constitute a detailed set of formal assertions and invariants that should hold at various steps in the FDLS algorithm. They are then inserted as programming assertions in the implementation of the FDLS algorithm in a production-strength high-level synthesis system. When turned on, the programming assertions (1) certify whether a specific run of the FDLS algorithm produced correct schedules and, (2) in the event of failure, help discover and isolate errors in the FDLS implementation.
Naren Narasimhan, Elena Teica, Rajesh Radhakrishnan, Sriram Govindarajan, Ranga Vemuri
ICCD5
1998 Automatic data path abstraction for verification of large scale designs
abstract
The state space explosion problem is a hurdle in the acceptance of model checking as a viable tool for verification of large-scale designs. Abstractions may be used to simplify designs, while preserving target verification properties. We propose a simple methodology for abstracting away portions of the data path, thus rendering a large state-space model of the design amenable for verification using model checking. The spatial abstractions developed reduce the bit-width complexity of the designs while retaining the controllers intact. The methodology uses interval computation techniques to determine the bounds on the allowable range of values the data path resources can assume. The approach is embedded in a tool that performs automatic data path abstraction on a RTL specification of a design.
Viresh Paruthi, Nazanin Mansouri, Ranga Vemuri
ICCD3
1997 A constructive method for data path area estimation during high-level VLSI synthesis
abstract
In this paper we present a fast and computationally efficient deterministic method for estimating the area of a register transfer level datapath obtained during high level VLSI synthesis. The estimation makes use of a RT level netlist along with a pre-synthesized library of RT level components. The layout area is estimated using a quadratic programming based framework to get a quick module allocation and generating a topological floorplan which is then followed by heuristic algorithms for mapping RTL modules and their interconnections on a standard cell based layout design style. Experiments on a suite of benchmark examples show promising results with reliable accuracy.
Natesan Venkateswaran, Srinivas Katkoori, Dinesh Bhatia, Ranga Vemuri
ASP-DAC5
1997 Symbolic Evaluation of Performance Models for Tradeoff Visualization
abstract
Often during the design process, it is necessary to analyzethe effects of tradeoff among various performance attributesof the design.Visual representations in the formof plots, graphs, and tables are typically used.These visualizationscan be generated using different applications suchas MatLab®, Mathematica®, Excel®, and so forth.Inorder to use these tools, performance equations for the designmust be rendered in an equational form with only afew variables, usually 2-3.However, performance modelsusually consist of a large number of attributes and evaluzationprocedures, typically written using the full power of aprogramming or hardware description language.We introducea symbolic evaluation procedure to simplify performancemodels.Given partial data, such evaluation yields a residualperformance model which is simpler than the original model.Symbolic evaluation can be effectively used to obtain residualperformance equations in terms of the variables whosetradeoff visualization is of interest to the user.
Jeffrey Walrath, Ranga Vemuri
DAC2
1997 Dynamic Bounding of Successor Force Computations in the Force Directed List Scheduling Algorithms
abstract
The well known Force Directed List Scheduling (FDLS) Algorithm uses a rigorous priority function called the Force of an operation. The force of an operation is governed by two components, namely the self-force of an operation and its successors' forces. The successor force in turn is governed by the self-force of all the descendants of the operation. FDLS is computationally intensive in its force calculations. For data flow dominated designs, a major portion of the FDLS execution time is spent in the computation of successor forces. However in this paper we observe that it is not always necessary to compute successor forces till the last successor level. We have shown in this paper that there usually exists a stabilization point after which successor force computations would not affect the quality of the schedule produced. This paper presents a concept of stability to show that it is possible to dynamically bound the successor force calculations in FDLS, up to a certain level of descendants. We have measured the performance of FDLS for a suite of high level synthesis benchmarks. Results presented in the paper show considerable reduction in execution time for the same schedule quality. This would allow a high-level synthesis tool to perform better design space exploration.
Sriram Govindarajan, Ranga Vemuri
ICCD2
1996 Rapid Prototyping of Reconfigurable Coprocessors
abstract
We describe the process of hardware-software codesign of a JPEG-like still image compression system. The hardware components are targeted to execute on a reconfigurable hardware coprocessor which communicates with a host computer that executes all the software tasks. Central to our codesign methodology is the usage of software profiling, high-level estimation and synthesis tools. We describe the process of trade-off analysis and hardware task selection in detail. We present detailed experimental results gathered throughout the codesign process.
Naren Narasimhan, Vinoo Srinivasan, Madhavi Vootukuru, Jeffrey Walrath, Sriram Govindarajan, Ranga Vemuri
ASAP6
1996 Specification of Control Flow Properties for Verification of Synthesized VHDL Designs
Naren Narasimhan, Ranga Vemuri
FMCAD2
1996 Simulation based architectural power estimation for PLA-based controllers
abstract
We present an architectural power simulation technique for PLA-based controllers. The contributions of this work are (1) a simple but efficient power characterization of PLAs; and (2) a strategy for developing a simulatable power model from the input description. Node Switching Capacitance (NSC) of a sub-component (such as AND plane) in a PLA is the average capacitance switched by a node in the sub-component, when the node undergoes a power consuming transition (0/spl rarr/1). Power characterization involves extracting NSC equations for different sub-components as a function of input size, output size and number of terms. Prototype PLAs whose are employed to derive NSC equations for a given technology. The input description is modified for power simulation by adding NSC equations with dependent variables instanced to the controller's parameters. For a given input sequence, the modified VHDL description is simulated to estimate the total power consumption. Experimental Results are obtained with average estimation error of 10.48% with a minimum error of 0.19% and maximum error of 21.90%.
Srinivas Katkoori, Ranga Vemuri
ISLPED2
1995 High level profiling based low power synthesis technique
abstract
We present a profiling based technique for power estimation. This technique is implemented in the PDSS (Profile Driven Synthesis System) for the synthesis of low power designs. Initially, each module in the module library is characterized for the average switching capacitance per input vector. The input description is simulated using user-specified set of input vectors to collect the profile data for various operators and carriers. The profile data, in conjunction with the pre-characterized module library is used to estimate the total capacitance switched by each of the valid schedules produced by the PDSS scheduler. A valid schedule is one which satisfies other constants such as area and delay. The schedule with the least switching capacitance estimate is further synthesized to the layout level. Results show an average deviation of 12% compared with the actual switching capacitance values at the layout level.
Srinivas Katkoori, Nand Kumar, Ranga Vemuri
ICCD3
1995 Generation of design verification tests from behavioral VHDL programs using path enumeration and constraint programming
abstract
A method for generation of design verification tests from behavior-level VHDL programs is presented. The method generates stimuli to execute desired control-flow paths in the given VHDL program. This method is based on path enumeration, constraint generation and constraint solving techniques that have been traditionally used for software testing. Behavioral VHDL programs contain multiple communicating processes, signal assignment statements, and wait statements which are not found in traditional software programming languages. Our model of constraint generation is specifically developed for VHDL programs with such constructs. Control-flow paths for which design verification tests are desired are specified through certain annotations attached to the control statements in the VHDL programs. These annotations are used to enumerate the desired paths. Each enumerated path is translated into a set of mathematical constraints corresponding to the statements in the path. Methods for generating constraint variables corresponding to various types of carriers in VHDL and for mapping various VHDL statements into mathematical relationships among these constraint variables are developed. These methods treat spatial and temporal incarnations of VHDL carriers as unique constraint variables thereby preserving the semantics of the behavioral VHDL programs. Constraints are generated in the constraint programming language CLP(R) and are solved using the CLP(R) system. A solution to the set of constraints so generated yields a design verification test sequence which can be applied for executing the corresponding control path when the design is simulated. If no solution exists, then it implies that the corresponding path can never be executed. Experimental studies pertaining to the quality of path coverage and fault coverage of the verification tests are presented.>
Ranga Vemuri, R. Kalyanaraman
IEEE Trans. Very Large Scale Integr. Syst.1
1993 Performance Specification Using Attributed Grammars
abstract
Article Performance specification using attributed grammars Share on Authors: Ram Mandayam View Profile , Ranga Vemuri View Profile Authors Info & Claims DAC '93: Proceedings of the 30th international Design Automation ConferenceJuly 1993 Pages 661–667https://doi.org/10.1145/157485.165085Online:01 July 1993Publication History 5citation151DownloadsMetricsTotal Citations5Total Downloads151Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ram Mandayam, Ranga Vemuri
DAC2
1993 Experiences in Functional Validation of a High Level Synthesis System
abstract
The goal OJ functional validation of a high-level synthesis system is to a$sert, with a reasonable degree of confidence, that the layouta generated by the high-level synthesis system correctly implement the ~pecified behavior.This paper presents a systematic approach to functional validation baaed on the analysis of specification language constructs, de8ign example formation to cover combinations of constructs, ted-bench generation and automated te8t re8u!t comparison.we have 8ucce88.fu![y applied thi8 approach in validating a high-level 8ynthe8i8 8 stem, called 33 DSS, which accepta specifications stated in VI-I L. This effort Te8u[ted in the development of a functional validation suite corssiding of 29 design ezample.s. 1
Ranga Vemuri, Paddy Mamtora, Praveen Sinha, Nand Kumar, Jayanta Roy, Raghu Vutukuru
DAC1
1992 Distributed Design-Space Exploration for High-Level Synthesis Systems
Rajiv Dutta, Jayanta Roy, Ranga Vemuri
DAC3
1991 Genetic synthesis: performance-driven logic synthesis using genetic evolution
abstract
The authors present a system for constraint-directed synthesis of logic circuits. The heart of the system is a genetic synthesis algorithm capable of searching the design space in the presence of user-specified constraints on performance attributes such as the area, critical path length, number of gates, wires and any other attribute which can be procedurally described. Three experimental results are presented and discussed.>
Ram Vemuri, Ranga Vemuri
Great Lakes Symposium on VLSI2
1978 Continuing education in microprocessors: Use of software simulators
abstract
This paper describes one school's experience in dealing with the continuing education needs of the community in which it is located. Specifically, the experience relates to the teaching of a course on Microprocessors and their applications.
Ranga Vemuri, J. V. Cornacchio
COMPSAC1