Deshanand P. Singh

dblp:74/3215 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
2since 2021 · last 2024
0009-0003-4968-4343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
Electronic design automation · 30% Cloud and datacenter computing · 30% Performance modeling and evaluation · 15%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
datacenter workloads
0.812024
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision · DAC 2024
Cloud and datacenter computing › inference serving
DNN serving
0.812024
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision · DAC 2024
Performance modeling and evaluation
workload characterization
0.812024
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision · DAC 2024
Electronic design automation
high-level synthesis
0.532015
High-Level Design Tools for Floating Point FPGAs · FPGA 2015
Profile-guided floating- to fixed-point conversion for hybrid FPGA-processor applications · ACM Trans. Archit. Code Optim. 2013
Harnessing the power of FPGAs using altera's OpenCL compiler · FPGA 2013
Electronic design automation
logic synthesis
0.342011
Line-level incremental resynthesis techniques for FPGAs · FPGA 2011
FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
FPGA technology mapping: a study of optimality · DAC 2005
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN inference
0.212024
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision · DAC 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.212024
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision · DAC 2024
Electronic design automation › high-level synthesis › arithmetic-level optimization
floating-point to fixed-point conversion
0.212013
Profile-guided floating- to fixed-point conversion for hybrid FPGA-processor applications · ACM Trans. Archit. Code Optim. 2013
Reconfigurable computing and FPGAs
FPGA accelerator
0.212013
Profile-guided floating- to fixed-point conversion for hybrid FPGA-processor applications · ACM Trans. Archit. Code Optim. 2013
Electronic design automation › logic synthesis
technology mapping
0.122007
FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
FPGA technology mapping: a study of optimality · DAC 2005
Reconfigurable computing and FPGAs
FPGA reliability
0.112010
A comprehensive approach to modeling, characterizing and optimizing for metastability in FPGAs · FPGA 2010
Integrated circuit design
metastability
0.112010
A comprehensive approach to modeling, characterizing and optimizing for metastability in FPGAs · FPGA 2010
Reconfigurable computing and FPGAs
FPGA physical design
0.122005
Incremental retiming for FPGA physical synthesis · DAC 2005
Integrated retiming and placement for field programmable gate arrays · FPGA 2002
Reconfigurable computing and FPGAs
FPGA architecture
0.122007
FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
The case for registered routing switches in field programmable gate arrays · FPGA 2001
Reconfigurable computing and FPGAs › FPGA architecture
configurable logic block
0.112007
FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation › design automation tools
FPGA CAD
0.122011
Line-level incremental resynthesis techniques for FPGAs · FPGA 2011
A comprehensive approach to modeling, characterizing and optimizing for metastability in FPGAs · FPGA 2010
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools
0.112015
High-Level Design Tools for Floating Point FPGAs · FPGA 2015
Electronic design automation › logic synthesis › sequential circuit optimization
retiming
0.122005
Incremental retiming for FPGA physical synthesis · DAC 2005
Integrated retiming and placement for field programmable gate arrays · FPGA 2002
Electronic design automation › logic synthesis › technology mapping
FPGA technology mapping
0.112005
FPGA technology mapping: a study of optimality · DAC 2005
Electronic design automation
physical design
0.112005
Incremental retiming for FPGA physical synthesis · DAC 2005
Program analysis › dynamic analysis › profiling
value profiling
0.012013
Profile-guided floating- to fixed-point conversion for hybrid FPGA-processor applications · ACM Trans. Archit. Code Optim. 2013
Parallel and multicore computing
parallel programming models
0.012013
Harnessing the power of FPGAs using altera's OpenCL compiler · FPGA 2013
Electronic design automation › physical design
timing optimization
0.012011
Line-level incremental resynthesis techniques for FPGAs · FPGA 2011
Electronic design automation › physical design › timing optimization
FPGA timing optimization
0.012002
Constrained clock shifting for field programmable gate arrays · FPGA 2002
Electronic design automation › physical design
placement
0.012002
Integrated retiming and placement for field programmable gate arrays · FPGA 2002
Reconfigurable computing and FPGAs
FPGA routing architecture
0.012001
The case for registered routing switches in field programmable gate arrays · FPGA 2001
Reconfigurable computing and FPGAs › FPGA architecture
FPGA clock network
0.012002
Constrained clock shifting for field programmable gate arrays · FPGA 2002

Methods — techniques the papers use, named apart from their topics

end-to-end latency and throughput measurement · 0.8empirical performance analysis · 0.8design space exploration · 0.3LLVM-based profiling · 0.3boolean satisfiability · 0.1HDL differencing · 0.1circuit-level simulation · 0.1board measurement · 0.1linear-time retiming · 0.1incremental clustering · 0.0
YearPublicationVenuePosition
2024 PEARL: Enabling Portable, Productive, and High-Performance Deep Reinforcement Learning using Heterogeneous Platforms
abstract
Deep Reinforcement Learning (DRL) is vital in various AI applications. DRL algorithms comprise diverse compute kernels, which may not be simultaneously optimized using a homogeneous architecture. However, even with available heterogeneous architectures, optimizing DRL performance remains a challenge due to the complexity of hardware and programming models employed in modern data centers. To address this, we introduce PEARL, a toolkit for composing parallel DRL systems on heterogeneous platforms consisting of general-purpose processors (CPUs) and accelerators (GPUs, FPGAs). Our innovations include: 1. A general training protocol agnostic of the underlying hardware, enabling portable implementations across various platforms. 2. Incorporation of DRL-specific optimizations on runtime scheduling and resource allocation, facilitating parallelized training and enhancing the overall system performance. 3. Automatic optimization of DRL task-to-device assignments through throughput estimation. 4. High-level API for productive development using the toolkit. We showcase our toolkit through experimentation with two widely used DRL algorithms, DQN and DDPG, on two diverse heterogeneous platforms. The generated implementations outperform state-of-the-art libraries for CPU-GPU platforms by up to 2.2× throughput improvements, and 2.4× higher performance portability across platforms.
Yuan Meng 0001, Michael Kinsner, Deshanand P. Singh, Mahesh A. Iyer, Viktor Prasanna 0001
CF3
2024 Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
abstract
Deep neural network (DNN) inference has become an important part of many data-center workloads. This has prompted focused efforts to design ever-faster deep learning accelerators such as GPUs and TPUs. However, an end-to-end DNN-based vision application contains more than just DNN inference, including input decompression, resizing, sampling, normalization, and data transfer. In this paper, we perform a thorough evaluation of computer vision inference requests performed on a throughput-optimized serving system. We quantify the performance impact of server overheads such as data movement, preprocessing, and message brokers between two DNNs producing outputs at different rates. Our empirical analysis encompasses many computer vision tasks including image classification, segmentation, detection, depth-estimation, and more complex processing pipelines with multiple DNNs. Our results consistently demonstrate that end-to-end application performance can easily be dominated by data processing and data movement functions (up to 56% of end-to-end latency in a medium-sized image, and ~ 80% impact on system throughput in a large image), even though these functions have been conventionally overlooked in deep learning system design. Our work identifies important performance bottlenecks in different application scenarios, achieves 2.25× better throughput compared to prior work, and paves the way for more holistic deep learning system design.
Ahmed F. AbouElhamayed, Susanne Balle, Deshanand P. Singh, Mohamed S. Abdelfattah
DAC3
2015 High-Level Design Tools for Floating Point FPGAs
abstract
This tutorial describes tools for efficiently implementing floating point applications on FPGAs. We present both the SDK for OpenCL and DSP Builder Advanced Blockset and show that they can be effectively used to implement many floating point applications. The methods for optimizing application performance are also described.
Deshanand P. Singh, Bogdan Pasca 0001, Tomasz S. Czajkowski
FPGA1
2013 Fractal video compression in OpenCL: An evaluation of CPUs, GPUs, and FPGAs as acceleration platforms
abstract
Fractal compression is an efficient technique for image and video encoding that uses the concept of self-referential codes. Although offering compression quality that matches or exceeds traditional techniques with a simpler and faster decoding process, fractal techniques have not gained widespread acceptance due to the computationally intensive nature of its encoding algorithm. In this paper, we present a real-time implementation of a fractal compression algorithm in OpenCL [1]. We show how the algorithm can be efficiently implemented in OpenCL and optimized for multi-CPUs, GPUs, and FPGAs. We demonstrate that the core computation implemented on the FPGA through OpenCL is 3× faster than a high-end GPU and 114× faster than a multi-core CPU, with significant power advantages. We also compare to a hand coded FPGA implementation to showcase the effectiveness of an OpenCL-to-FPGA compilation tool.
Doris Chen, Deshanand P. Singh
ASP-DAC2
2013 Harnessing the power of FPGAs using altera's OpenCL compiler
abstract
In recent years, Field-Programmable Gate Arrays have become extremely powerful computational platforms that can efficiently solve many complex problems. The most modern FPGAs comprise effectively millions of programmable elements, signal processing elements and high-speed interfaces, all of which are necessary to deliver a complete solution. The power of FPGAs is unlocked via low-level programming languages such as VHDL and Verilog, which allow designers to explicitly specify the behavior of each programmable element. While these languages provide a means to create highly efficient logic circuits, they are akin to "assembly language" programming for modern processors. This is a serious limiting factor for both productivity and the adoption of FPGAs on a wider scale. In this talk, we use the OpenCL language to explore techniques that allow us to program FPGAs at a level of abstraction closer to traditional software-centric approaches. OpenCL is an industry standard parallel language based on 'C' that offers numerous advantages that enable designers to take full advantage of the capabilities offered by FPGAs, while providing a high-level design entry language that is familiar to a wide range of programmers.
Deshanand P. Singh, Tomasz S. Czajkowski, Andrew C. Ling
FPGA1
2013 Profile-guided floating- to fixed-point conversion for hybrid FPGA-processor applications
abstract
The key to enabling widespread use of FPGAs for algorithm acceleration is to allow programmers to create efficient designs without the time-consuming hardware design process. Programmers are used to developing scientific and mathematical algorithms in high-level languages (C/C++) using floating point data types. Although easy to implement, the dynamic range provided by floating point is not necessary in many applications; more efficient implementations can be realized using fixed point arithmetic. While this topic has been studied previously [Han et al. 2006; Olson et al. 1999; Gaffar et al. 2004; Aamodt and Chow 1999], the degree of full automation has always been lacking. We present a novel design flow for cases where FPGAs are used to offload computations from a microprocessor. Our LLVM-based algorithm inserts value profiling code into an unmodified C/C++ application to guide its automatic conversion to fixed point. This allows for fast and accurate design space exploration on a host microprocessor before any accelerators are mapped to the FPGA. Through experimental results, we demonstrate that fixed-point conversion can yield resource savings of up to 2x--3x reductions. Embedded RAM usage is minimized, and 13%--22% higherFmaxthan the original floating-point implementation is observed. In a case study, we show that 17% reduction in logic and 24% reduction in register usage can be realized by using our algorithm in conjunction with a High-Level Synthesis (HLS) tool.
Doris Chen, Deshanand P. Singh
ACM Trans. Archit. Code Optim.2
2012 Invited paper: Using OpenCL to evaluate the efficiency of CPUS, GPUS and FPGAS for information filtering
abstract
The FPGA can be a tremendously efficient computational fabric for many applications. In particular, the performance to power ratios of FPGA make them attractive solutions to solve the problem of data centers that are constrained largely by power and cooling costs. However, the complexity of the FPGA design flow requires the programmer to understand cycle-accurate details of how data is moved and transformed through the fabric. In this paper, we explore techniques that allow programmers to efficiently use FPGAs at a level of abstraction that is closer to traditional software-centric approaches by using the emerging parallel language, OpenCL. Although the field of high level synthesis has evolved greatly in the last few decades, several fundamental parts were missing from the complete software abstraction of the FPGA. These include standard and portable methods of describing HW/SW codesign, memory hierarchy, data movement and control of parallelism. We believe that OpenCL addresses all of these issues and allows for highly efficient description of FPGA designs with a higher level of abstraction. We demonstrate this premise by examining the performance of a document filtering algorithm, implemented in OpenCL and automatically compiled to a Stratix IV 530 FPGA. We show that our implementation achieves 5.5× and 5.25× better performance per watt ratios than GPU and CPU implementations, respectively.
Doris Chen, Deshanand P. Singh
FPL2
2012 From opencl to high-performance hardware on FPGAS
abstract
We present an OpenCL compilation framework to generate high-performance hardware for FPGAs. For an OpenCL application comprising a host program and a set of kernels, it compiles the host program, generates Verilog HDL for each kernel, compiles the circuit using Altera Complete Design Suite 12.0, and downloads the compiled design onto an FPGA.We can then run the application by executing the host program on a Windows(tm)-based machine, which communicates with kernels on an FPGA using a PCIe interface. We implement four applications on an Altera Stratix IV and present the throughput and area results for each application. We show that we can achieve a clock frequency in excess of 160MHz on our benchmarks, and that OpenCL computing paradigm is a viable design entry method for high-performance computing applications on FPGAs.
Tomasz S. Czajkowski, Utku Aydonat, Dmitry Denisenko, John Freeman, Michael Kinsner, David Neto, Peter Yiannacouras, Deshanand P. Singh
FPL9
2011 Line-level incremental resynthesis techniques for FPGAs
abstract
FPGA logic density is roughly doubling at every process generation. Consequently, it is becoming increasingly challenging for FPGA CAD tools to keep up with the growing complexities of high-speed designs while keeping CAD run-times reasonable. In this paper, we present a novel incremental resynthesis tool called Line-Level Incremental reSynthesis (LLIS), integrated within an industrial tool suite, that addresses the problems of timing closure as well as CAD runtime (patent pending). We describe a general framework that can incrementally reuse results from a previous compile based on automatic differencing of HDL changes. We show that it is possible to reduce synthesis runtime by 6.5x for common HDL changes. As compared with complete resynthesis, we preserve known good timing solutions more than 82% of the time. This represents a 3X improvement vs. non-incremental techniques.
Doris Chen, Deshanand P. Singh
FPGA2
2010 A comprehensive approach to modeling, characterizing and optimizing for metastability in FPGAs
abstract
Metastability is a phenomenon that can cause system failures in digital circuits. It may occur whenever signals are being transmitted across asynchronous or unrelated clock domains. The impact of metastability is increasing as process geometries shrink and supply voltages drop faster than transistor Vts. FPGA technologies are significantly affected since leading edge FPGAs are amongst the first devices to adopt the most recent process nodes. In this paper, we present a comprehensive suite of techniques for modeling, characterizing and optimizing metastability effects in FPGAs. We first discuss a theoretical model of metastability, and verify the predictions using both circuit level simulations and board measurements. Next we show how designers have traditionally dealt with metastability problems and contrast that with the automatic CAD algorithms described in this paper that both analyze and optimize metastability-related issues. Through our detailed experimental results, we show that we can improve the metastability characteristics of a large suite of industrial benchmarks by an average of 268,000 times with our optimization techniques.
Doris Chen, Deshanand P. Singh, Jeffrey Chromczak, David M. Lewis, Ryan Fung, David Neto, Vaughn Betz
FPGA2
2010 Parallelizing FPGA Technology Mapping Using Graphics Processing Units (GPUs)
abstract
GPUs are becoming an increasingly attractive option for obtaining performance speedups for data-parallel applications. FPGA technology mapping is an algorithm that is heavily data parallel; however, it has many features that make it unattractive to implement on a GPU. The algorithm uses data in irregular ways since it is a graph-based algorithm. In addition, it makes heavy use of constructs like recursion which is not supported by GPU hardware. In this paper, we take a state-of-the-art FPGA technology mapping algorithm within Berkeley's ABC package and attempt to parallelize it on a GPU. We show that runtime gains of 3.1× are achievable while maintaining identical quality as demonstrated by running these netlists through Altera's Quartus II place-and-route tool.
Doris Chen, Deshanand P. Singh
FPL2
2007 Incremental placement for structured ASICs using the transportation problem
abstract
While physically driven synthesis techniques have proven to be an effective method to meet tight timing constraints required by a design, the incremental placement step during physically driven synthesis has emerged as the primary bottleneck. As a solution, this paper introduces a scalable incremental placement algorithm based upon the well known transportation problem. This method has an average speedup of 2× and a 30% reduction in memory usage when compared against a commercial incremental placer without any impact on area or speed of the final placed circuit. Furthermore, this method is scalable for structured ASICs.
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003
VLSI-SoC2
2007 An area-efficient timing closure technique for FPGAs using Shannon's expansion
Deshanand P. Singh, Stephen Brown 0003
Integr.1
2007 FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability
abstract
This paper presents a field-programmable gate array (FPGA) logic synthesis technique based upon Boolean satisfiability. This paper shows how to map any Boolean function into an arbitrary programmable logic block (PLB) architecture without any custom decomposition techniques. The authors illustrate several useful applications of this technique by showing how this technique can be used for architecture evaluation and area optimization. When evaluating the FPGA architecture, the authors focus on the basic building block of the FPGA, which they refer to as PLB. In order to illustrate the flexibility of their evaluation framework, several unrelated PLB architectures are evaluated in an automated fashion. Furthermore, the authors show that using their technique is able to reduce FPGA resource usage by 27% on average in common subcircuits found in digital design.
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Predicting Interconnect Delay for Physical Synthesis in a FPGA CAD Flow
abstract
This paper studies the prediction of interconnect delay in an industrial setting. Industrial circuits and two industrial field-programmable gate-array (FPGA) architectures were used in this paper. We show that there is a large amount of inherent randomness in a state-of-the-art FPGA placement algorithm. Thus, it is impossible to predict interconnect delay with a high degree of accuracy. Furthermore, we show that a simple timing model can be used to predict some aspects of interconnect timing with just as much accuracy as predictions obtained by running the placement tool itself. Using this simple timing model in a two-phase timing driven physical synthesis flow can both improve quality of results and decrease runtime. Next, we present a metric for predicting the accuracy of our interconnect delay model and show how this metric can be used to reduce the runtime of a timing driven physical synthesis flow. Finally, we examine the benefits of using the simple timing model in a timing driven physical synthesis flow, and attempt to establish an upper bound on these possible gains, given the difficulty of interconnect delay prediction.
Valavan Manohararajah, Gordon R. Chiu, Deshanand P. Singh, Stephen Brown 0003
IEEE Trans. Very Large Scale Integr. Syst.3
2006 Mapping arbitrary logic functions into synchronous embedded memories for area reduction on FPGAs
abstract
This work describes a new mapping technique, RAM-MAP, that identifies parts of circuits that can be efficiently mapped into the synchronous embedded memories found on field programmable gate arrays (FPGAs). Previous techniques developed for mapping into asynchronous embedded memories cannot be used because modern FPGAs do not have asynchronous embedded memories. After technology mapping, an area-prediction cost function is used to guide the selection of logic cones to be placed in embedded memories. Extra logic is added to compensate for missing asynchronous functionality on the synchronous memories. Experiments conducted on Altera's Stratix device family indicate that this embedded memory mapping technique can provide an average area reduction of 6.2% and up to 32.5% on a large set of industrial designs. A small architecture change that increases the size of the FPGA fabric by 0.05% can increase the average area reduction to 14.1% and up to 59.1% on the same design set.
Gordon R. Chiu, Deshanand P. Singh, Valavan Manohararajah, Stephen Brown 0003
ICCAD2
2005 FPGA technology mapping: a study of optimality
abstract
This paper attempts to quantify the optimality of FPGA technology mapping algorithms. We develop an algorithm, based on Boolean satisfiability (SAT), that is able to map a small subcircuit into the smallest possible number of lookup tables (LUTs) needed to realize its functionality. We iteratively apply this technique to small portions of circuits that have already been technology mapped by the best available mapping algorithms for FPGAs. In many cases, the optimal mapping of the subcircuit uses fewer LUTs than is obtained by the technology mapping algorithm. We show that for some circuits the total area improvement can be up to 67%.
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003
DAC2
2005 Incremental retiming for FPGA physical synthesis
abstract
In this paper, we present a new linear-time retiming algorithm that produces near-optimal results. Our implementation is specically targeted at Altera's Stratix [1] FPGA-based designs, although the techniques described are general enough for any implementation medium. The algorithm is able to handle the architectural constraints of the target device, multiple timing constraints assigned by the user and implicit legality constraints. It ensures that register moves do not create asynchonous problems such as creating a glitch on a clock/reset signal.
Deshanand P. Singh, Valavan Manohararajah, Stephen Brown 0003
DAC1
2005 FPGA PLB Evaluation using Quantified Boolean Satisfiability
abstract
This paper describes a novel field programmable gate array (FPGA) logic synthesis technique which determines if a logic function can be implemented in a given programmable circuit and describes how this problem can be formalized and solved using quantified Boolean satisfiability. This technique is general enough to be applied to any type of logic function and programmable circuit; thus, it has many applications to FPGAs. The application demonstrated in this paper is FPGA PLB evaluation where their results show that this tool allows radical new features of FPGA logic blocks to be evaluated in a rigorous scientific way.
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003
FPL2
2005 Post-Placement BDD-Based Decomposition for FPGAs
abstract
This work explores the effect of adding a timing driven functional decomposition step to the traditional field programmable gate array (FPGA) CAD flow. Once placement has completed, alternative decompositions of the logic on the critical path are examined for potential delay improvements. The placed circuit is then modified to use the best decompositions found. Any placement illegalities introduced by the new decompositions are resolved by an incremental placement step. Experiments conducted on Altera's Stratix and Stratix II device families indicate that this functional decomposition technique can provide average performance improvements of 6.1% and 5.6% on a large set of industrial designs, respectively.
Valavan Manohararajah, Deshanand P. Singh, Stephen Brown 0003
FPL2
2005 FPGA Logic Synthesis Using Quantified Boolean Satisfiability
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003
SAT2
2004 The Quartus University Interface Program: enabling advanced FPGA research
abstract
FPGA researchers constantly face the challenge of determining whether their innovations will work in the real world. The Quartus University Interface Program (QUIP) allows the researcher to answer this important question by directly integrating research prototypes within one of the FPGA industry's leading CAD tool suites. This work describes the QUIP interface as well as research projects that are of significant interest to the FPGA industry.
Shawn Malhotra, Terry P. Borer, Deshanand P. Singh, Stephen Brown 0003
FPT3
2002 Integrated retiming and placement for field programmable gate arrays
abstract
Retiming is a synchronous circuit transformation that can optimize the delay of a synchronous circuit by moving registers across combinational circuit elements. The combinational structure remains unchanged and the observable behavior of the circuit is identical to the original.In this paper, we address the problem of applying retiming techniques to circuits implemented in Field Programmable Gate Arrays (FPGAs). FPGAs contain prefabricated and configurable routing elements that allow us to easily implement a variety of circuits. However this interconnect contributes greatly to the overall delay in the implemented circuit. If a circuit is retimed prior to the placement and routing phases of the CAD flow, then it has no information about the delays introduced by the configurable interconnect. Our fundamental experiment is to determine whether there are any gains in tightly coupling retiming and placement so that the retiming algorithm has some estimate of the routing delays.Specifically, we introduce a post-placement retiming algorithm that understands how to take advantage of FPGA architectural features. This retiming algorithm may introduce extra registers into the circuit. These new registers need to be placed in some location in the FPGA. Retiming register placement is accomplished by a novel incremental clustering and placement algorithm. The incremental algorithm builds upon the placement of the non-retimed circuit to intelligently sift in the newly-introduced registers.In addition, we explore making the placement algorithms "retiming aware." These placement algorithms try to place logic blocks in such a way that the subsequent retiming produces better speed results. These techniques include the identification of retiming-critical cycles during placement.Our experiments show that the integration of retiming with placement results in 19% better clock periods in comparison to the application of retiming before the place and route steps.
Deshanand P. Singh, Stephen Brown 0003
FPGA1
2002 Constrained clock shifting for field programmable gate arrays
abstract
Circuits implemented in FPGAs have delays that are dominated by its programmable interconnect. This interconnect provides the ability to implement arbitrary connections. However, it contains both highly capacitive and resistive elements. The delay encountered by any connection depends strongly on the number of interconnect elements used to route the connection. These delays are only completely known after the place and route phase of the CAD flow. We propose the use of Clock Shifting optimization techniques to improve the clock frequency as a post place and route step.Clock Shifting Optimization is a technique first formalized in [4]. It is a cycle-stealing algorithm that allows one to reduce the critical path delay of a synchronous circuit by shifting the clock signals at each register. This technique allows late arriving signals to be sampled at a later point in time by intentionally introducing a skew on the clock input of the sampling register. Typical FPGAs contain a number of special purpose global clock networks that distribute clock signals to every register in the chip. Unused global clock lines in FPGAs can be used to distribute a finite set of clock skews to the entire circuit. We propose an efficient integer programming method to find the optimal circuit improvement for a finite set of clock skews. This technique is modified to consider inherent uncertainties present in the timing models. The uncertainty controls the aggressiveness of the optimizations as we must take great care in ensuring functionality for any range of possible timing characteristics.Our results confirm intuition that more aggressive speed optimizations can be performed as timing models become more accurate. We also show that providing 4 skewed versions of the nominal clock signal results in the best delay--area tradeoff. This result is evocative as it may suggest future FPGA architectures that contain greater numbers of global clock lines, as we tradeoff gains in speed for greater power requirements from increased clock network flexibility.
Deshanand P. Singh, Stephen Brown 0003
FPGA1
2002 Incremental placement for layout driven optimizations on FPGAs
abstract
This paper presents an algorithm to update the placement of logic elements when given an incremental netlist change. Specifically, these algorithms are targeted to incrementally place logic elements created by layout-driven circuit restructuring techniques. The incremental placement engine assumes that the restructuring algorithms provide a list of new logic elements along with preferred locations for each of these new elements. It then tries to shift non-critical logic elements in the original placement out of the way to satisfy the preferred location requests. Our algorithm considers modern FPGA architectures with clustered logic blocksthat have numerous architectural constraints. Experiments indicate that our technique produces results of extremely highquality.
Deshanand P. Singh, Stephen Brown 0003
ICCAD1
2001 The case for registered routing switches in field programmable gate arrays
abstract
FPGAs are characterized by a programmable interconnect that contains highly resistive and capacitive elements. While the configurable structure of the interconnect allows for the implementation of arbitrary circuits, it has also become a significant bottleneck for high-speed circuits. Even if there are only a few signal paths that run along long stretches of interconnect, it is these paths that may determine the maximum operating frequency of the circuit.
Deshanand P. Singh, Stephen Brown 0003
FPGA1