Dinesh Bhatia

dblp:03/4333 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Electronic design automation · 58% Performance modeling and evaluation · 26% Reconfigurable computing and FPGAs · 16%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › benchmarking
benchmark dataset
0.412020
MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows · FPGA 2020
Electronic design automation
high-level synthesis
0.412020
MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows · FPGA 2020
Electronic design automation
machine learning for EDA
0.412020
MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows · FPGA 2020
Performance modeling and evaluation
performance prediction
0.412020
MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows · FPGA 2020
Electronic design automation
physical design
0.452016
Floorplanning of Partially Reconfigurable Design on Heterogeneous FPGA (Abstract Only) · FPGA 2016
Interconnect estimation for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
A priori wirelength and interconnect estimation based on circuit characteristic · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design
floorplanning
0.212016
Floorplanning of Partially Reconfigurable Design on Heterogeneous FPGA (Abstract Only) · FPGA 2016
Reconfigurable computing and FPGAs › dynamic reconfiguration
partial reconfiguration
0.212016
Floorplanning of Partially Reconfigurable Design on Heterogeneous FPGA (Abstract Only) · FPGA 2016
Electronic design automation › physical design › routing
routability
0.122006
Interconnect estimation for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
On metrics for comparing routability estimation methods for FPGAs · DAC 2002
Electronic design automation
simulated annealing
0.112016
Floorplanning of Partially Reconfigurable Design on Heterogeneous FPGA (Abstract Only) · FPGA 2016
Reconfigurable computing and FPGAs
FPGA physical design
0.112006
Interconnect estimation for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Reconfigurable computing and FPGAs
FPGA architecture
0.112005
A priori wirelength and interconnect estimation based on circuit characteristic · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design › placement
wirelength estimation
0.112005
A priori wirelength and interconnect estimation based on circuit characteristic · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › physical design › floorplanning
FPGA floorplanning
0.021999
A Methodology for Fast FPGA Floorplanning · FPGA 1999
Performance Driven Floorplanning for FPGA Based Designs · FPGA 1997
Electronic design automation › physical design › placement › constructive placement
clustering-based placement
0.011999
A Methodology for Fast FPGA Floorplanning · FPGA 1999
Electronic design automation › high-level synthesis › scheduling
dataflow graph scheduling
0.011999
Temporal Partitioning and Scheduling Data Flow Graphs for Reconfigurable Computers · IEEE Trans. Computers 1999
Electronic design automation › physical design
placement
0.011999
A Methodology for Fast FPGA Floorplanning · FPGA 1999
Reconfigurable computing and FPGAs › FPGA partitioning
temporal partitioning
0.011999
Temporal Partitioning and Scheduling Data Flow Graphs for Reconfigurable Computers · IEEE Trans. Computers 1999
Electronic design automation › physical design › floorplanning
timing-driven floorplanning
0.011997
Performance Driven Floorplanning for FPGA Based Designs · FPGA 1997
Electronic design automation › hardware verification and test › design for testability
built-in self-test
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation
hardware verification and test
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › hardware verification and test › test generation
pseudoexhaustive testing
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › hardware verification and test
sequential circuit testing
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › hardware verification and test
test generation
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › physical design
routing
0.012002
On metrics for comparing routability estimation methods for FPGAs · DAC 2002
Reconfigurable computing and FPGAs › FPGA-based emulation
logic emulation
0.011999
Temporal Partitioning and Scheduling Data Flow Graphs for Reconfigurable Computers · IEEE Trans. Computers 1999
Electronic design automation
logic synthesis
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › logic synthesis › sequential circuit optimization
retiming
0.011995
Pseudo-exhaustive built-in TPG for sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995

Methods — techniques the papers use, named apart from their topics

statistical analysis · 0.4machine learning · 0.4white space detection · 0.2simulated annealing · 0.2priority-based sorting · 0.2routing flexibility · 0.1router characterization · 0.1structural circuit analysis · 0.1detailed router comparison · 0.0benchmark circuits · 0.0
YearPublicationVenuePosition
2023 Enhancing Student Engagement in Engineering and Education Through Virtual Reality: A Survey-Based Analysis
abstract
This paper investigates the impact of virtual reality (VR) on student engagement in engineering education and their potential in enhancing the student learning experience through technology-led learning. VR technology has shown to improve student engagement in different educational sectors. As a result, this study evaluated the efficacy of VR technology as an immersive learning tool focusing on engineering education. The survey collects data from a diverse sample of students who experienced the use of VR-based flight simulator as a part of their continuous assessment, providing valuable insights. Conditional results indicate that by using VR headsets 70% of students reported improved learning outcomes for the module, while 100% students agreed that VR technology offered a more immersive learning experience. The survey results indicate and emphasize the potential benefits that integrating VR headsets into engineering education could bring to the student learning experience through enhanced student engagement and the acquisition of practical skills in simulated immersive environments. The paper also makes recommendations for further research and implementation of VR technology in other fields of STEM education.
Dinesh Bhatia, Henrik Hesse
TENCON1
2023 Machine learning based fast and accurate High Level Synthesis design space exploration: From graph to synthesis
Pingakshya Goswami, Benjamin Carrión Schäfer, Dinesh Bhatia
Integr.3
2020 MLSBench: A Synthesizable Dataset of HLS Designs to Support ML Based Design Flows
abstract
With the advent of Machine Learning (ML), predictive EDA tools are becoming the next hot topic of research in the EDA community, and researchers are working on ML-based tools to predict the performance of the EDA tool. As the designs become complex, there is a need to start the design using higher levels of abstraction, such as High-Level Synthesis (HLS) tools in FPGA and SoC design flows. Quick prediction of performance-related parameters of the final design after the C-synthesis stage, can help in rapid design closure. Even though multiple papers exist in the domain of post routing performance prediction of HLS tools, there are no standard benchmarks available to compare the performance and accuracy of the predictive models. In this paper, we have presented MLSBench, a collection of around 5000 synthesizable designs written in C and C++. We provide a methodology to generate designs with various variations from a single design, which creates a potential for creating newer designs and enlarging the database in the future. This is followed by analysis, and validating the generated designs are indeed different. This allows designers to create generalized machine-learning-based models that are not overfitted to a small dataset. We also perform statistical analysis for measuring the design diversity by synthesizing them using Xilinx-Vivado HLS for Zynq 7000 device series.
Pingakshya Goswami, Masoud Shahshahani, Dinesh Bhatia
FPGA3
2020 Empirical greedy machine-based automatic liver segmentation in CT images
abstract
Segmentation of the liver from 3D computed tomography volumes plays a significant role in trajectory development for computer‐assisted interventional surgery for the liver disease. Despite a lot of studies, liver segmentation remains a challenging task due to the lack of clear edges on most liver boundaries coupled with high variability of both anatomical and intensity patterns. In addition, there is a problem with the segmentation of the left portal vein, in which the size of this vein prominently estimates the liver tumour area. The empirical greedy machine is proposed to make the precise, automated segmentation of the liver as well as the left portal vein. In which the empirical robust nature trains the features of the liver proficiently thereby segmenting the liver from other organs without the omission of adjacent organs and liver lobe region. Hence this proposed method can achieve one of the highest accuracies compared to other segmentation methods and the performance is calculated using several parameters such as volumetric overlap error, relative absolute volume difference (RVD), average symmetric absolute surface distance (ASD), root mean square surface distance, maximum symmetric ASD.
Gajendra Kumar Mourya, Dinesh Bhatia, Akash Handique
IET Image Process.2
2017 Portable impedance measurement device for sweat based glucose detection
abstract
The future of disease diagnostics and health care wearables lies in the development of low-cost sensors that can detect minute traces of pathogens or antigens from body fluids. Developments in nanotechnology and biomedical research have already shown us that a nanosensor can be specifically tailored to detect a specific biomolecule. These sensors would allow patients to run point of care diagnostic tests, thereby saving time and cost of running clinical tests and can give early stage disease diagnosis and help physicians to provide personalized treatment. This work involves the development of a configurable electronic sensor platform that will interface with these sensors. The device is tested by quantification of glucose from sweat using a nanosensor developed in the Biomedical Microdevices and Nanotechnology Lab in the University of Texas at Dallas. The platform can be easily configured to run Electrochemical Impedance Spectroscopy based detection test for other biomolecules by using sensor tailored for it.
Athul Asokan Thulasi, Dinesh Bhatia, Poras T. Balsara, Shalini Prasad
BSN2
2016 Floorplanning of Partially Reconfigurable Design on Heterogeneous FPGA (Abstract Only)
abstract
The floorplanning problem in FPGA has been a topic of research for more than a decade. Although the floorplanning problem has been thoroughly explored for homogeneous FPGAs, very less work has been done for heterogeneous FPGAs. In this paper, we have designed a floorplanner for partially reconfigurable design in heterogeneous FPGAs which takes into consideration the diversity of resources present inside the FPGA device and their locations. The floorplanner is based on fixed outline simulated annealing algorithm. We proposed a priority based sorting algorithm mimicking Olympic Medal Tally for initial floorplanning which is a preprocessing step for simulated annealing. Also, a White Space detection algorithm is proposed for efficient management of white space inside the FPGA device. We defined a cost function, which consists of weighed sum of wirelength, area and resource wastage, and minimized this cost function using simulated annealing. In this work, we also described a method to calculate two different types of resource wastage and suggested a method to reduce it. The performance of our floorplanner is evaluated using MCNC benchmarks on Xilinx Virtex 5 FPGA architecture. We have compared our proposed floorplanner with other results reported in the literature and observed substantial improvement in the overall wirelength as well as the execution time. Finally, we integrated our floorplanner with Xilinx PlanAhead to automate the floorplanning process. We generate the floorplan of a partially reconfigurable median filter, which consists of seven reconfigurable regions. When a comparison is made between the manually generated floorplan and the automatic floorplan generated by our tool, a significant improvement is observed in terms of parameters like total area occupied by the reconfigurable regions, frequency of the operation and total time required by PlanAhead to place and route the design.
Pingakshya Goswami, Dinesh Bhatia
FPGA2
2008 A dynamic temperature control simulation system for FPGAs
abstract
Rapid increases in transistor density, clock speeds and competition with custom ICs have escalated the demand for aggressive solutions to battle rising operating temperatures in programmable fabrics. In this work, we make several key contributions to temperature management in FPGAs. We develop a novel and robust simulation framework exploring adaptive techniques to reduce on chip temperatures in the reconfigurable core. We implement a thermal driven voltage scaling algorithm based on temperature and performance feedback. Our performance estimation model is an accurate empirical relation between delay, supply voltage and temperature with an average error of 9%. Our final results show significant temperature reductions of up to 13.37degC accompanied by the added benefit of power savings averaging 13.48%. Overheads are limited to an average reduction in worst case operating frequency of 10.78% and a voltage swing of 0.61V.
Shilpa Bhoj, Dinesh Bhatia
FPL2
2008 Early stage FPGA interconnect leakage power estimation
abstract
Increasing transistor densities, rising popularity in mobile applications and migration towards eco-friendly computing systems have made power dissipation a key FPGA design issue. To meet stringent budgets, system architects need accurate estimates of power distribution at various design stages. In this work, we make several key contributions to FPGA leakage power estimation. First, we develop an accurate and efficient model to estimate total interconnect leakage power at various design stages prior to routing. Our methods derive leakage power estimates based on predicted values of routing congestion and interconnect resource utilization. We then extend the model to accomodate complex segmented routing architectures and low leakage architectures. Finally we formulate relations to generate post place leakage power estimates of individual routing channels. Our models for overall leakage power estimation achieve average accuracy rates of 93% and 89% for uniform and segmented routing architectures respectively. Experimentation results also establish the accuracy of the channel level estimation models at 85% and 80% for uniform and segmented routing structures. Our models and techniques would help designers make informed decisions by providing information on the power consumption of the interconnect fabric well before routing. Additionally, the equations can be used for architectural explorations and embedded in power and thermal aware CAD tools.
Shilpa Bhoj, Dinesh Bhatia
ICCD2
2007 Pre-route Interconnect Capacitance and Power Estimation in FPGAs
abstract
The increase in functional complexity and performance requirements of reconfigurable fabrics has necessitated the early estimation of power distribution to explore power aware architectures and design techniques. In this work, we make several contributions to interconnect power estimation in FPGAs. First we present a probabilistic methodology that uses demand, a parameter reflecting the probable usage of a routing resource, to estimate interconnect capacitance and dynamic power dissipation. We then extend this model to estimate static power distribution of interconnect resources. Finally, the framework of our model also allows us to estimate spatial power distribution. Our results indicate accurate predictions with average errors of 7, 11 and 7 percent for capacitance, dynamic and static power respectively.
Shilpa Bhoj, Dinesh Bhatia
FPL2
2007 Thermal Modeling and Temperature Driven Placement for FPGAs
abstract
With rapid technology scaling and increased use of FPGAs in mobile and low power applications, the effect of temperature on power, performance and reliability has become a challenge for system designers. In this work we make two major contributions to thermal aware design in FPGAs. First, we formulate a resistive mesh model for thermal profiling of FPGAs. The model exploits the characteristic features of the reconfigurable fabric to produce an efficient and accurate thermal estimate. We then employ the thermal model in a novel temperature driven placement tool for FPGAs. Our placement tool achieves an average temperature reduction of 2.26degC and produces a more uniform thermal distribution. The impact on frequency and wirelength is small with an average degradation of 5.28 and 4.15 percent respectively.
Shilpa Bhoj, Dinesh Bhatia
ISCAS2
2006 A leakage aware design methodology for power-gated programmable architectures
abstract
One popular technique to deal with increasing leakage current in modern FPGAs is to group logic into sleeper cells and isolate them from power rails whenever they are inactive in time and space. However, such architectures demand novel methodology to effectively layout the design, so that maximum sleeper cells can be shut down. In this work, we present a slicing tree based simulated annealing methodology that grooms the various blocks of a design, so that they can be put into power-down mode either in temporal or spatial domain. Our technique nicely fits into a power-aware design flow targeting standby/portable applications where substantial portion of the design sits idle for a long period of time. We also propose a novel technique to include BRAMs into the presented methodology. Our experiments shows up to a maximum 43% of leakage savings in some benchmarks
Narayan Subramanian, Rajarshee P. Bharadwaj, Dinesh Bhatia
FPT3
2006 Interconnect estimation for FPGAs
abstract
Interconnect planning is becoming an important design issue for large field programmable gate array (FPGA)-based designs. One of the most important issues for planning interconnection is the ability to reliably predict the routing requirements of a given design. In this paper, a new methodology, called fast generic routability estimation for placed FPGA circuits (fGREP), for fast and reliable estimation of routing requirements for placed circuits on island-style FPGAs, is introduced. This method is based on newly derived detailed router characterizations that are introduced in this paper. It is observed that the router has a limited number of available routing elements to use and the number is proportional to the distance from a net's terminal. This is defined as the routing flexibility and an estimate for interconnect requirements is derived from it. This method is able to predict the distribution of interconnect requirements, with very fine granularity, across the entire device. The interconnect-distribution information is used to estimate congestion and total wirelength. Multiterminal nets are efficiently handled, without the need for net decomposition. This method is generic enough to enable its usage with any standard FPGA place-and-route design flow and for any island-style FPGA architecture. The method is also applicable to application-specific integrated circuit (ASIC) design flows. Experimental results on a large set of standard benchmark examples show that the estimates obtained here closely match with the detailed routing results of the state-of-the-art router PathFinder , as implemented in the well-known FPGA physical design suite VPR.
PariVallal Kannan, Dinesh Bhatia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Exploiting temporal idleness to reduce leakage power in programmable architectures
abstract
One of the biggest challenges that programmable devices like FPGAs are facing in ultra deep sub-micron regime is the exponential rise in leakage power consumption. As technology shrinks below 90nm, a new design paradigm has to evolve to tackle the issue of leakage power consumption. In this work we focus on a new design methodology for reducing leakage power by exploiting temporal locality in designs and accordingly group them into. clusters that can be switched on and off. We propose a Power State Controller based method, which controls the switching of the clusters from one state to another. We show our technique using Data Flow Graphs where temporal locality can be effectively explored. Our results show that substantial leakage savings can be achieved if temporal idleness of designs can be exploited effectively.
Rajarshee P. Bharadwaj, Rajan Konar, Poras T. Balsara, Dinesh Bhatia
ASP-DAC4
2005 Timing Aware Interconnect Prediction Models for FPGAs
abstract
A-priori interconnect prediction is the estimation of routing requirements of a design without even performing placement. All current a-priori models assume that the optimization objective during placement and routing is to minimize wirelength and congestion. However, design performance is a very important goal to all designers and interconnects requirements for a design that are optimized for timing have not been studied. We present an empirical methodology that takes c, a user defined input that trades off routability and timing, and estimate a-priori some important characteristics of the design. This tunable model will be very useful for design under uncertainties or preliminary feasibility study at system levels. We show how the length of source-sink pairs is important and how its distribution can be predicted. Our results show very accurate prediction for circuits that are placed and routed with VPR.
Shankar Balachandran, Dinesh Bhatia
FPL2
2005 FPGA Architecture for Standby Power Management
Rajarshee P. Bharadwaj, Rajan Konar, Dinesh Bhatia, Poras T. Balsara
FPT3
2005 A priori wirelength and interconnect estimation based on circuit characteristic
abstract
Interconnect prediction is very important for early feasibility studies in modern design flows. Most of the current interconnect estimation techniques estimate either the average or the total wirelength and some qualitative measure of routing demand for circuits. A priori techniques estimate these parameters without actually performing circuit placement. We propose a new a priori interconnect and wirelength estimation methodology for island style field programmable gate arrays (FPGAs). For a given design, we estimate bounding box semiperimeter wirelengths of all nets for an optimized placement and the minimum number of tracks per channel required for successful routing on an FPGA device. We analyze the structural characteristics of circuits and limitations posed by the FPGA architecture to derive a consistent model for wirelength and routing demand estimation. We identify reconvergences present in a circuit as an important global circuit characteristic in wirelength prediction. Our overall results show that we have an average error of 11.6% w.r.t. semiperimeter wirelength measured from the optimized layout using VPR. Also, the number of routing tracks per channel is predicted with an average error of 13.2% of the detailed routing results from VPR.
Shankar Balachandran, Dinesh Bhatia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 On metrics for comparing interconnect estimation methods for FPGAs
abstract
Interconnect management is a critical design issue for large field-programmable gate arrays (FPGA) based designs. One of the most important issues for planning interconnection is the ability to reliably and efficiently predict the interconnect requirements of a given design on a given FPGA architecture. Many interconnect estimation methods have been reported so far and the estimation problem is also under active research. From a CAD tool deployment point of view, comparing these estimation methods is very difficult because of the different reporting methods used by the authors. We make an argument for and propose a new uniform reporting metric, based on comparing the estimates with the results of an actual detailed router on both local and global levels. We then compare some of the well known and promising interconnect estimation methods using our new metric on a large number of benchmark circuits.
PariVallal Kannan, Shankar Balachandran, Dinesh Bhatia
IEEE Trans. Very Large Scale Integr. Syst.3
2003 FPGA based EBCOT architecture for JPEG 2000
abstract
In this paper a high speed FPGA based implementation of Embedded Block Coding with Optimized Truncation (EBCOT) algorithm used in JPEG 2000 has been proposed and implemented. The context formation engine used in EBCOT is analyzed and an architecture based on parallel processing of the three coding passes is proposed. The architecture is coded in VHDL and the design is targeted to Xilinx Virtex II FPGA family. When implemented on a XC2V1000 device, the design performs at 56 MHz after place and route. Simulation results show that the design can process a 512/spl times/512 image in less than 0.03 seconds and the processing time is reduced by more than 75% compared to sample based implementation and by more than 34% compared to the best architecture known.
Manjunath Gangadhar, Dinesh Bhatia
FPT2
2003 Interconnect Estimation for FPGAs under Timing Driven Domains
abstract
Interconnect planning is fast becoming an important design issue for large FPGA based designs. The fundamental requirement for interconnect planning is the ability to estimate the routing requirements of a given design. Many estimation methods for uniform island-style FPGA architectures have been reported. However, no estimation method targets the problem of estimating the interconnect requirements for timing driven physical design. Most estimation methods assume minimum cost routing, which underestimates the interconnect resource requirements when timing is of main concern. Timing driven physical design typically involves minimum delay routing which demands extra interconnects as compared to minimum-cost routing. We propose a new method to estimate the interconnect requirements of placed FPGA circuits, under timing driven domains. We compare our estimates with the detailed routing results produced by standard routing tools in the VPR [V. Betz et al., (1997)] design suite.
PariVallal Kannan, Dinesh Bhatia
ICCD2
2002 On metrics for comparing routability estimation methods for FPGAs
abstract
Interconnect management is a critical design issue for large FPGA based designs. One of the most important issues for planning interconnection is the ability to accurately and efficiently predict the routability of a given design on a given FPGA architecture. The recently proposed routability estimation procedure, fGREP [6], produced estimates within 3 to 4% of an actual detailed router. Other known routability estimation methods include RISA [5], Lois's [7] method and Rent's rule based methods [1] [11] [9]. Comparing these methods has been difficult because of the different reporting methods used by the authors. We propose a uniform reporting metric based on comparing the estimates produced with the results of an actual detailed router on both local and global levels. We compare all the above methods using our reporting metric on a large number of benchmark circuits and show that the enhanced fGREP method produces tight estimates that outperform most other techniques.
PariVallal Kannan, Shankar Balachandran, Dinesh Bhatia
DAC3
2002 Rapid and Reliable Routability Estimation for FPGAs
PariVallal Kannan, Shankar Balachandran, Dinesh Bhatia
FPL3
2001 Tightly Integrated Placement and Routing for FPGAs
PariVallal Kannan, Dinesh Bhatia
FPL2
2001 fGREP - Fast Generic Routing Demand Estimation for Placed FPGA Circuits
PariVallal Kannan, Shankar Balachandran, Dinesh Bhatia
FPL3
2000 Bounds, designs and layouts for multi-terminal FPIC architectures
Dinesh Bhatia, James Haralambides
Integr.1
2000 Resource requirements and layouts for field programmable interconnection chips
abstract
Field-programmable interconnection chips (FPIC's) provide the capability of realizing user programmable interconnection for any desired permutation. Such an interconnection is very much desired for supporting rapid prototyping of hardware systems and for providing programmable communication networks for parallel and distributed computing. An FPIC should realize any possible permutation of input to output pins via a set of programmable switches. In this paper, we show that any such architecture requires a minimum of /spl Omega/(n log n) switches, where /spl Omega/ is the number of I/O pins. The result stems from an analysis of the underlying permutation network. In addition, for networks of bounded degree d, we prove an /spl Omega/(log/sub d-1/ n) bound on the routing delay (maximum length of routing paths for specific I/O permutations) and an /spl Omega/(n log/sub d-1/ n) bound on the average utilization of programmable switches used by the FPIC to implement a specific permutation. For the same type of networks, we prove an /spl Omega/(n log/sub d-1/ n) bound on the number of nodes of the network. Furthermore, we design efficient architectures for FPIC's offering a wide variety of routing delays, high average programmable resource utilization, and O(n/sup 2/)-area two-layer layouts. The proposed structures are called hybrid Benes-Crossbar (HBC) architectures and clearly exhibit a tradeoff between performance (routing delay utilization) and area of the layout.
Dinesh Bhatia, James Haralambides
IEEE Trans. Very Large Scale Integr. Syst.1
1999 A Methodology for Fast FPGA Floorplanning
abstract
Floorplanning is an important problem in FPGA circuit mapping.As FPGA capacity grows, new innovative approaches will be required for eficiently mapping circuits to FPGAs.In this paper we present a macro basedjoorplanning methodology suitable for mapping large circuits to large, high density FPGAs.Our method uses clustering techniques to combine macros into clusters, and then uses a tabu search based approach to place clusters while enhancing both circuit routability and performance.Our method is capable of handling both hard (axed size and shape) macros and soft (fixed size and variable shape) macros.We demonstrate our methodology on several macro based circuit designs and compare the execution speed and qualiw of results with commercially available CAE tools.Our approach shows a dramatic speedup in execution time without any negative impact on quality.
John Marty Emmert, Dinesh Bhatia
FPGA2
1999 Temporal Partitioning and Scheduling Data Flow Graphs for Reconfigurable Computers
abstract
FPGA-based configurable computing machines are evolving rapidly. They offer the ability to deliver very high performance at a fraction of the cost when compared to supercomputers. The first generation of configurable computers (those with multiple FPGAs connected using a specific interconnect) used statically reconfigurable FPGAs. On these configurable computers, computations are performed by partitioning an entire task into spatially interconnected subtasks. Such configurable computers are used in logic emulation systems and for functional verification of hardware. In general, configurable computers provide the ability to reconfigure rapidly to any desired custom form. Hence, the available resources can be reused effectively to cut down the hardware costs and also improve the performance. In this paper, we introduce the concept of temporal partitioning to partition a task into temporally interconnected subtasks. Specifically, we present algorithms for temporal partitioning and scheduling data flow graphs for configurable computers. We are given a configurable computing unit (RPU) with a logic capacity of S/sub RPU/ and a computational task represented by an acyclic data flow graph G=(V, E). Computations with logic area requirements that exceed S/sub RPU/ cannot be completely mapped on a configurable computer (using traditional spatial mapping techniques). However, a temporal partitioning of the data flow graph followed by proper scheduling can facilitate the configurable computer based execution. Temporal partitioning of the data flow graph is a k-way partitioning of G=(V, E) such that each partitioned segment will not exceed S/sub RPU/ in its logic requirement. Scheduling assigns an execution order to the partitioned segments so as to ensure proper execution. Thus, for each segment in {s/sub 1/,s/sub 2/,...,s/sub k/}, scheduling assigns a unique ordering S/sub i/-j,1/spl les/i/spl les/k,1/spl les/j/spl les/k, such that the computation would execute in proper sequential order as defined by the flow graph G=(V, E).
Karthikeya M. Gajjala Purna, Dinesh Bhatia
IEEE Trans. Computers2
1998 Temporal Partitioning and Scheduling for Reconfigurable Computing
abstract
FPGA based custom computing machine applications have grown tremendously. Reconfigurable FPGAs incur very less reconfiguration times and also have the ability to reconfigure partially. They provide avenues to reuse the hardware resources at runtime, thus decreasing the hardware costs. In this paper, we present algorithms for temporal partitioning of applications into small size segments (under the area constraints), and scheduling of segments to ensure proper execution by satisfying the data dependencies among the segments. Our investigation concentrates on applications that are also directed acyclic graphs (DAGs). We have implemented the algorithms and have produced mappings of real applications on reconfigurable hardware.
Karthikeya M. Gajjala Purna, Dinesh Bhatia
FCCM2
1998 Partitioning in time: a paradigm for reconfigurable computing
abstract
In recent years, we have witnessed the rapid growth of reconfigurable computers. The first generation of reconfigurable computers consists of multiple FPGAs interconnected in a network. The computations are performed by partitioning an entire task into spatially interconnected sub-tasks. FPGAs used in the reconfigurable computers are programmed only once during the runtime of an executing application. FPGAs have the ability to reconfigure rapidly to any desired custom form. Reusing the FPGA resources during the application runtime can yield cost effective solutions for reconfigurable computing. Such runtime reconfiguration of FPGAs requires an efficient framework for the analysis and synthesis of the application and tools that handle the runtime reconfiguration. In this paper, we introduce the concept of temporal partitioning, to partition a task into temporally interconnected sub-tasks. We present algorithms and methodologies to analyze an application, and techniques to reuse the programmable hardware during the runtime of the application. Our approach has been successfully tested on read applications and has proven to be cost effective.
Karthikeya M. Gajjala Purna, Dinesh Bhatia
ICCD2
1998 Clock-skew constrained placement for row based designs
abstract
In this paper we address the problem of placement of standard cells under the constraints of minimizing the clock-skew. We propose a quadratic programming based methodology for placement that not only results in an area and timing wise good placement but also a supporting zero-skew clock routing tree. Under the clock-skew constraints, our method produces significant reduction in the cost of zero-skew clock routing tree. During placement, we are able to obtain significant speed-up due to variable reduction and constraint modification.
Natesan Venkateswaran, Dinesh Bhatia
ICCD2
1997 A constructive method for data path area estimation during high-level VLSI synthesis
abstract
In this paper we present a fast and computationally efficient deterministic method for estimating the area of a register transfer level datapath obtained during high level VLSI synthesis. The estimation makes use of a RT level netlist along with a pre-synthesized library of RT level components. The layout area is estimated using a quadratic programming based framework to get a quick module allocation and generating a topological floorplan which is then followed by heuristic algorithms for mapping RTL modules and their interconnections on a standard cell based layout design style. Experiments on a suite of benchmark examples show promising results with reliable accuracy.
Natesan Venkateswaran, Srinivas Katkoori, Dinesh Bhatia, Ranga Vemuri
ASP-DAC4
1997 Performance Driven Floorplanning for FPGA Based Designs
abstract
Increasing design densities on large FPGAs and greater demand for performance, has calledfor special purpose tools like floorplanner, performance driven router, and more. In this paper we present a floorplanning based design mapping solution that is capable of mapping macro cell based designs as well as hierarchicaldesigns on FPGAs. The mapping solution has been tested extensively on a large collection of designs. We not only outperform state of the art CAE tools from industry in terms of execution time but also achieve much better performance in terms of timing. These methods are especially suitable for mapping designs on very large FPGAs.
Jianzhong Shi, Dinesh Bhatia
FPGA2
1997 Partitioning Under Timing and Area Constraints
abstract
Circuit partitioning is a very extensively studied problem. In this paper we formulate the problem as a nonlinear program (NLP). The NLP is solved for the objective of minimum cutset size under the constraints of timing. Our proposed methodology easily extends to multiple constraints that are very dominant in the design of large scale VLSI Systems. The NLP is solved using the commercial LP/NLP solver MINOS. We have done extensive testing using large scale RT level benchmarks and have shown that our methods can be used for exploring the design space for obtaining constraint satisfying system designs. We also provide extensions for solving system design problems where a choice between multiple technologies, packaging components, performance, cost, yield, and more can be the constraints for design related decisions.
Gregory Tumbush, Dinesh Bhatia
ICCD2
1996 Multiway Partitioner for High Performance FPGA Based Board Architecture
abstract
Field-programmable gate array based board architectures are becoming fairly common for rapid prototyping and custom computing. In order to map large designs on multiple FPGA based boards, the design has to be partitioned into two or more segments. In this paper we describe the architecture, constraints, and a solution to the area and pin constrained partitioning problem. Our effort is directed towards partitioning "huge" designs in relatively small amount of time, thus giving the designer a capability to explore many mapping solutions. The board level architecture is based on multi-chip modules, where each MCM consists of three Xilinz 4025 FPGAs and a dual ported one mega byte SRAM. In its smallest configuration the board can map 300,000 gate size designs.
Vijayanand Sankarasubramanian, Dinesh Bhatia
ICCD2
1995 Pseudo-exhaustive built-in TPG for sequential circuits
abstract
We address the issue of pseudo-exhaustive test pattern generation (TPG) for the built-in self-test (BIST) of sequential circuits. Let d be the sequential depth, and w be the input dependency limit. We use an LFSR/SR Test Pattern Generator and a small additional hardware overhead to automatically generate d/spl middot/2/sup w/ test patterns to test the circuit pseudo exhaustively or, alternatively, pseudo-randomly with less hardware overhead and extremely high fault coverage. Our scheme uses novel retiming algorithms and transforms the circuit to an equivalent (for test purposes) one by scanning a subset of flip-flops for breaking its cyclic structure, bounding the sequential depth, forcing the input dependency limit, balancing the circuit, and maintaining the clock period. We present the first polynomial time algorithm to bound the sequential depth of a circuit by retiming with minimum number of flip-flops and subject to a clock period bound. We also give a retiming-based polynomial time algorithm to balance a circuit by inserting a minimum number of bypass delay cells. Experimental results on the ISCAS'89 benchmarks indicate that our method outperforms a previously proposed approach, which not only does not provide for on-chip test pattern generation but also requires O(q/spl middot/f/spl middot/2/sup w/) test patterns, where q is the total number of primary or pseudo-primary outputs in the circuit and f is the total number of flip-flops.>
Dimitrios Kagaris, Spyros Tragoudas, Dinesh Bhatia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1994 Mathematical model for routability analysis of FPGAs
abstract
We have developed a mathematical model for estimating the probability of routing on an electrically programmable logic cell array (LGA). Using the model the variation of the routability or the probability of routing with the average span of the nets, and flexibility of programming resources is determined. The results obtained from the model have been verified experimentally. As expected, the routability also increases when the number of programming elements is increased. We also show a three dimensional relationship between routability, placement, and programming flexibility. Our results show a strong relationship between module placement and the routability for an LCA.>
Dinesh Bhatia, Amit Chowdhary, Spyros Tragoudas
Great Lakes Symposium on VLSI1
1994 Generalized segmented channel routing
abstract
This paper presents the first efficient solution to the generalized detailed routing problem in segmented channels for row-based FPGAs. A generalized detailed routing allows routing of each connection using an arbitrary number of tracks, i.e. doglegs are allowed. This approach is different from the normally followed method where each connection is routed on a single straight track. We present a router that performs generalized segmented channel routing using a greedy approach to route channels. It uses effective data-structures and pruning heuristics to keep down the time and memory requirements of the router.>
V. Shankar, Dinesh Bhatia
Great Lakes Symposium on VLSI2
1993 Pseudoexhaustive BIST for Sequential Circuits
abstract
We present a method that can be used to test a sequential circuit pseudoexhaustively or almost pseudoexhaustively using LFSR/SRs as ATPGs with d-2/sup w/ test patterns, where d is the sequential depth and w is the input dependency limit. Our approach is based on the following techniques: (1) Use of LFSR/SRs as ATPGs (2) Rearrangement of the flip-flops of the circuit by retiming so that the hardware overhead for breaking all cycles and bounding the sequential depth is minimized. (3) Introduction of bypass storage cells (BSCs) so that no combinational element in the circuit has input dependence greater than a user-defined constant w. (4) Introduction of bypass delay cells (BDCs) so that the graph becomes more easily balanced or approximately balanced. Comparative experimental results indicate that our method behaves better than full-scan. It also outperforms a previous approach which, not only does not provide for on-chip TPG, but also requires O(q-f-2/sup 2/) test patterns, where q is the total number of primary or pseudoprimary outputs in the circuit and f is the total number of flip-flops.>
Dimitrios Kagaris, Spyros Tragoudas, Dinesh Bhatia
ICCD3
1992 Improved Algorithms for Routing on Two-Dimensional Grids
Dinesh Bhatia, Frank Thomson Leighton, Fillia Makedon, Carolyn Haibt Norton
WG1