EDBT 2026 Demo / reviewers in the wild / expert
Stephen Brown 0003
dblp:b/StephenDeanBrown · also Stephen Dean Brown
· DBLP profile ↗
72ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
28 papers |
Electronic design automation · 63% Reconfigurable computing and FPGAs · 27% Energy-efficient computing · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 30 heaviest of 52, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.7 | 4 | 2016 | A Survey and Evaluation of FPGA High-Level Synthesis Tools · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 High-level synthesis with LegUp: a crash course for users and researchers · FPGA 2013 Impact of FPGA architecture on resource sharing in high-level synthesis · FPGA 2012 |
Electronic design automation
logic synthesis |
0.5 | 9 | 2008 | Scalable Synthesis and Clustering Techniques Using Decision Diagrams · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Functionally Linear Decomposition and Synthesis of Logic Circuits for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Functionally linear decomposition and synthesis of logic circuits for FPGAs · DAC 2008 |
Reconfigurable computing and FPGAs
FPGA-accelerated simulation |
0.4 | 1 | 2020 | Using OpenCL to Enable Software-like Development of an FPGA-Accelerated Biophotonic Cancer Treatment Simulator · FPGA 2020 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.3 | 2 | 2016 | A Survey and Evaluation of FPGA High-Level Synthesis Tools · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 Towards automated ECOs in FPGAs · FPGA 2009 |
Reconfigurable computing and FPGAs
FPGA high-level synthesis |
0.2 | 1 | 2016 | A Survey and Evaluation of FPGA High-Level Synthesis Tools · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016 |
Electronic design automation › logic synthesis
technology mapping |
0.2 | 6 | 2007 | FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Heuristics for Area Minimization in LUT-Based FPGA Technology Mapping · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 FPGA technology mapping: a study of optimality · DAC 2005 |
Electronic design automation › physical design
engineering change order |
0.2 | 2 | 2011 | Toward Automated ECOs in FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Towards automated ECOs in FPGAs · FPGA 2009 |
Reconfigurable computing and FPGAs
FPGA design flow |
0.2 | 2 | 2011 | Toward Automated ECOs in FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Scalable Synthesis and Clustering Techniques Using Decision Diagrams · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Electronic design automation › logic synthesis › technology mapping
FPGA technology mapping |
0.2 | 3 | 2008 | Functionally Linear Decomposition and Synthesis of Logic Circuits for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Heuristics for Area Minimization in LUT-Based FPGA Technology Mapping · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 FPGA technology mapping: a study of optimality · DAC 2005 |
Electronic design automation
physical design |
0.2 | 3 | 2011 | Toward Automated ECOs in FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Incremental retiming for FPGA physical synthesis · DAC 2005 A detailed router for field-programmable gate arrays · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992 |
Energy-efficient computing › low-power design
power optimization |
0.2 | 2 | 2010 | Decomposition-Based Vectorless Toggle Rate Computation for FPGA Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 Using Negative Edge Triggered FFs to Reduce Glitching Power in FPGA Circuits · DAC 2007 |
Electronic design automation › high-level synthesis › hardware compilation
c-to-hardware compilation |
0.2 | 1 | 2013 | High-level synthesis with LegUp: a crash course for users and researchers · FPGA 2013 |
Electronic design automation › hardware/software co-design
hardware/software partitioning |
0.2 | 1 | 2013 | High-level synthesis with LegUp: a crash course for users and researchers · FPGA 2013 |
Distributed systems
resource sharing |
0.1 | 1 | 2012 | Impact of FPGA architecture on resource sharing in high-level synthesis · FPGA 2012 |
Electronic design automation › logic synthesis › logic optimization
resynthesis |
0.1 | 1 | 2011 | Toward Automated ECOs in FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 |
Electronic design automation › power estimation
FPGA power estimation |
0.1 | 1 | 2010 | Decomposition-Based Vectorless Toggle Rate Computation for FPGA Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Reconfigurable computing and FPGAs
FPGA architecture |
0.1 | 4 | 2007 | FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Hybrid FPGA Architecture · FPGA 1996 The case for registered routing switches in field programmable gate arrays · FPGA 2001 |
Reconfigurable computing and FPGAs
FPGA physical design |
0.1 | 2 | 2005 | Incremental retiming for FPGA physical synthesis · DAC 2005 Integrated retiming and placement for field programmable gate arrays · FPGA 2002 |
Electronic design automation › logic synthesis
BDD-based synthesis |
0.1 | 1 | 2008 | Scalable Synthesis and Clustering Techniques Using Decision Diagrams · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Electronic design automation
clustering |
0.1 | 1 | 2008 | Scalable Synthesis and Clustering Techniques Using Decision Diagrams · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Electronic design automation › logic synthesis › technology mapping › FPGA technology mapping
lookup table mapping |
0.1 | 1 | 2008 | Functionally Linear Decomposition and Synthesis of Logic Circuits for FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Reconfigurable computing and FPGAs › FPGA architecture
configurable logic block |
0.1 | 1 | 2007 | FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean Satisfiability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Energy-efficient computing › dynamic power reduction
glitch reduction |
0.1 | 1 | 2007 | Using Negative Edge Triggered FFs to Reduce Glitching Power in FPGA Circuits · DAC 2007 |
Electronic design automation › logic synthesis › sequential circuit optimization
retiming |
0.1 | 2 | 2005 | Incremental retiming for FPGA physical synthesis · DAC 2005 Integrated retiming and placement for field programmable gate arrays · FPGA 2002 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.0 | 1 | 2013 | High-level synthesis with LegUp: a crash course for users and researchers · FPGA 2013 |
Electronic design automation › physical design
circuit clustering |
0.0 | 1 | 2003 | Recursive circuit clustering for minimum delay and area · FPGA 2003 |
Electronic design automation › physical design › timing optimization
FPGA timing optimization |
0.0 | 1 | 2002 | Constrained clock shifting for field programmable gate arrays · FPGA 2002 |
Electronic design automation › physical design
placement |
0.0 | 1 | 2002 | Integrated retiming and placement for field programmable gate arrays · FPGA 2002 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2010 | Decomposition-Based Vectorless Toggle Rate Computation for FPGA Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Reconfigurable computing and FPGAs
FPGA routing architecture |
0.0 | 1 | 2001 | The case for registered routing switches in field programmable gate arrays · FPGA 2001 |
Methods — techniques the papers use, named apart from their topics
OpenCL · 0.9iterative numerical methods · 0.4iterative numerical method · 0.4benchmarking methodology · 0.2high-level synthesis · 0.2functional linear decomposition · 0.2binary decision diagram · 0.2sharing cost/benefit analysis · 0.1binding phase optimization · 0.1boolean satisfiability · 0.1c-to-hardware compilation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Leveraging Fine-grained Structured Sparsity for CNN Inference on Systolic Array ArchitecturesabstractThe high computational complexity of convolutional neural networks (CNNs) has motivated many studies of accelerating CNN inference on field-programmable gate arrays (FPGAs). Among these, designs that feature systolic arrays can effectively leverage the parallelism in CNNs while acheiving good placement and routing quality. Weight sparsity – the presence of zeros in CNN weights – can further reduce the number of necessary multiply-accumulate (MAC) operations in CNNs, but has yet resulted in performance gain on systolic arrays. In this work, we propose a novel fine-grained structured weight sparsity pattern, showcase a processing element (PE) design that leverages this sparsity pattern, and develop a systolic array CNN inference accelerator that targets an Intel Arria 10 GX1150 FPGA. When evaluated on ResNet-50 and VGG-16 that are trained and pruned on the ImageNet dataset, our accelerator achieves 2.26 TOPs/s and 1.21 TOPs/s, respectively, on the MAC operations, while keeping the top-l accuracy degradation within 5%. These results translate to $2.86\times$ and $1.75\times$ speed-up compared to a dense systolic array baseline. Linqiao Liu, Stephen Brown 0003 |
FPL | 2 |
| 2020 | Using OpenCL to Enable Software-like Development of an FPGA-Accelerated Biophotonic Cancer Treatment SimulatorabstractThe simulation of light propagation through tissues is important for medical applications, such as photodynamic therapy (PDT) for cancer treatment. To optimize PDT an inverse problem, which works backwards from a desired distribution of light to the parameters that caused it, must be solved. These problems have no closed-form solution and therefore must be solved numerically using an iterative method. This involves running many forward light propagation simulations which is time-consuming and computationally intensive. Tanner Young-Schultz, Lothar Lilge, Stephen Brown 0003, Vaughn Betz |
FPGA | 3 |
| 2017 | FISH: Linux system calls for FPGA acceleratorsabstractThis, paper presents the FISH (FPGA-Initiated Software-Handled) framework which allows FPGA accelerators to make system calls to the Linux operating system in CPU-FPGA systems. A special FISH Linux kernel module running on the CPU provides a system call interface for FPGA accelerators, much like the ABI which exists for software programs. We provide a proof-of-concept implementation of this framework running on the Intel Cyclone V SoC device, and show that an FPGA accelerator can seamlessly make system calls as if it were the host program. We see the FISH framework being especially useful for high-level synthesis (HLS) by making it possible to synthesize software code that contains system calls. Kevin Nam, Blair Fort, Stephen Brown 0003 |
FPL | 3 |
| 2017 | From Pthreads to Multicore Hardware Systems in LegUp High-Level Synthesis for FPGAsabstractIn the last decade, processor speeds have remained fairly stagnant, and to improve performance further, the industry started to increase the number of processor cores. The use of specialized hardware, such as field-programmable gate arrays (FPGAs), has also been on the rise. The traditional design methodology for FPGAs, however, requires hardware knowledge, which makes the platform inaccessible to software engineers. High-level synthesis (HLS) tools aim to resolve this issue by allowing software design methodologies to be used for FPGAs. However, HLS remains difficult to use for many software engineers, as there are tasks, such as system integration, which is still mostly a manual process. Consequently, creating a multicore hardware system on an FPGA is not feasible for most software engineers. To this end, we provide an HLS framework, which can automatically generate a multicore hardware system from software. We provide support for POSIX threads, which can be compiled to concurrently executing hardware cores that can be used in a processor-accelerator hybrid system, or in a hardware-only system without a processor. With this, we show that we can create multicore FPGA systems that can provide significant benefits in performance and energy-efficiency compared with hardware executing sequentially, and software executing on MIPS/ARM/x86 processors. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A unified software approach to specify pipeline and spatial parallelism in FPGA hardwareabstractHigh-level synthesis (HLS) is increasingly becoming a mainstream design methodology for FPGAs. Whereas its previous applications were mostly limited to research and simple designs, it is now being used to tape-out real-world chips in production [1]. Advances in compiler and HLS research continue to improve the quality of HLS-generated hardware. Despite this, the ease-of-use of HLS tools remains a hurdle to its broad uptake, particularly by engineers without hardware skills. To this end, we propose using a well-known software technique to infer streaming parallel hardware in HLS. Specifically, we use the producer-consumer pattern, commonly used in multi-threaded programming, to infer the generation of hardware that can exploit both pipeline and spatial parallelism on FPGAs. Our proposed methodology allows one to create a design in software, using only standard software methodologies, that cannot only synthesize to streaming hardware, but also model the generated hardware more accurately than existing solutions from other state-of-the-art C-based HLS tools. We use four different real-world benchmarks to illustrate the use of our methodology, and how it can create circuits that are either pipelined, or pipelined and replicated, all from software. For comparison, we also use a commercial HLS tool to synthesize one of the benchmarks, and show that our methodology can produce competitive results to that of the commercial tool. Jongsok Choi, Ruolong Lian, Stephen Brown 0003, Jason Helge Anderson |
ASAP | 3 |
| 2016 | A Survey and Evaluation of FPGA High-Level Synthesis ToolsabstractHigh-level synthesis (HLS) is increasingly popular for the design of high-performance and energy-efficient heterogeneous systems, shortening time-to-market and addressing today’s system complexity. HLS allows designers to work at a higher-level of abstraction by using a software program to specify the hardware functionality. Additionally, HLS is particularly interesting for designing field-programmable gate array circuits, where hardware implementations can be easily refined and replaced in the target device. Recent years have seen much activity in the HLS research community, with a plethora of HLS tool offerings, from both industry and academia. All these tools may have different input languages, perform different internal optimizations, and produce results of different quality, even for the very same input description. Hence, it is challenging to compare their performance and understand which is the best for the hardware to be implemented. We present a comprehensive analysis of recent HLS tools, as well as overview the areas of active interest in the HLS research community. We also present a first-published methodology to evaluate different HLS tools. We use our methodology to compare one commercial and three academic tools on a common set ofCbenchmarks, aiming at performing an in-depth evaluation in terms of performance and the use of resources. Razvan Nane, Vlad Mihai Sima, Christian Pilato, Jongsok Choi, Blair Fort, Andrew Canis, Yu Ting Chen, Hsuan Hsiao, Stephen Brown 0003, Fabrizio Ferrandi, Jason Helge Anderson, Koen Bertels |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2015 | Resource and memory management techniques for the high-level synthesis of software threads into parallel FPGA hardwareabstractRecent work has proposed the high-level synthesis of parallel software programs (specified using Pthreads or OpenMP) into concurrently operating parallel hardware modules [6]. In this paper, we describe resource and memory management techniques for improving performance and area of hardware generated by such software thread synthesis. One direction investigated pertains to how modules in the HLS-generated parallel hardware should connect to one another: 1) with a nested topology, or 2) with a flat topology. In the nested topology, hardware modules are created in a hierarchical manner: modules are instantiated inside within modules that use them. Conversely, the flat topology instantiates all hardware modules at the same level of hierarchy. For the flat topology, we describe a system generator that automatically generates the required interconnect between all hardware modules, as well as flexibly shares or replicates functions, functional units, and memories. We also explore methods to reduce memory contention among hardware units that operate in parallel, by investigating three different memory architectures which use: 1) a global memory controller, 2) local memories, and 3) shared-local memories. Local and shared-local memories are dedicated RAM blocks for a single or a set of hardware modules, and help to increase memory bandwidth by allowing concurrent memory accesses. We also consider memory replication to localize memories in hardware modules, and convert small memories to registers to further improve performance and memory usage. Finally, we describe implementing locks and barriers in HLS hardware: synchronization constructs used in parallel programming. We show that with our resource and memory management techniques, we can improve the geomean performance, area, and area-delay product of parallel HLS-generated hardware up to 41.6%, 38.3%, and 63.3%, respectively, for a set of 15 benchmarks. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
FPT | 2 |
| 2015 | The Effect of Compiler Optimizations on High-Level Synthesis-Generated HardwareabstractWe consider the impact of compiler optimizations on the quality of high-level synthesis (HLS)-generated field-programmable gate array (FPGA) hardware. Using an HLS tool implemented within the state-of-the-art LLVM compiler, we study the effect of compiler optimizations on the hardware metrics of circuit area, execution cycles, FMax , and wall-clock time. We evaluate 56 different compiler optimizations implemented within LLVM and show that some optimizations significantly affect hardware quality. Moreover, we show that hardware quality is also affected by some optimization parameter values, as well as the order in which optimizations are applied. We then present a new HLS-directed approach to compiler optimizations, wherein we execute partial HLS and profiling at intermittent points in the optimization process and use the results to judiciously undo the impact of optimization passes predicted to be damaging to the generated hardware quality. Results show that our approach produces circuits with 16% better speed performance, on average, versus using the standard -O3 optimization level. Qijing Huang 0001, Ruolong Lian, Andrew Canis, Jongsok Choi, Ryan Xi, Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2014 | Automating the Design of Processor/Accelerator Embedded Systems with LegUp High-Level SynthesisabstractLegUp [1] is an open-source high-level synthesis (HLS) tool that accepts a C program as input and automatically synthesizes it into a hybrid system. The hybrid system comprises an embedded processor and custom accelerators that realize user-designated compute-intensive parts of the program with improved throughput and energy efficiency. In this paper, we overview the LegUp framework and describe several recent developments: 1) support for an embedded ARM processor, as is available on Altera's recently released SoC FPGA, 2) HLS support for software parallelization schemes -- pthreads and OpenMP, 3) enhancements to LegUp's core HLS algorithms that raise the quality of the auto-generated hardware, and, 4) a preliminary debugging and verification framework providing C source-level debugging of HLS hardware. Since its first release in 2011, LegUp has been downloaded over 1000 times by groups around the world, providing a powerful platform for new research in high-level synthesis algorithms and embedded systems design. Blair Fort, Andrew Canis, Jongsok Choi, Nazanin Calagar, Ruolong Lian, Stefan Hadjis, Yu Ting Chen, Mathew Hall, Bain Syrowik, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
EUC | 11 |
| 2014 | Source-level debugging for FPGA high-level synthesisabstractWe describe a source-level debugging framework for FPGA high-level synthesis (HLS) that offers gdb-like step, break, and data inspection functionality for an HLS-generated hardware circuit. With the proposed framework, the user can inspect the values of logic signals in the hardware from the C source code perspective. The logic signal values come from one of two sources: 1) a logic simulation of the RTL, or 2) an actual execution of the hardware on an FPGA. In addition to the software-like ecosystem for FPGA HLS debugging, the framework provides the user with insight on the RTL produced by the HLS tool for each C statement, and permits concurrent hardware and software debugging to discover the first point at which any logic signal in the hardware mismatches with its corresponding variable in software. Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
FPL | 2 |
| 2014 | Modulo SDC scheduling with recurrence minimization in high-level synthesisabstractLoop pipelining is a high-level synthesis scheduling technique that overlaps loop iterations to achieve higher performance. However, industrial designs often have resource constraints and other constraints imposed by cross-iteration dependencies. The interaction between multiple constraints can pose a challenge for HLS modulo scheduling algorithms, which, if not handled properly can lead to a loop pipeline schedule that fails to achieve the minimum possible initiation interval. We propose a novel modulo scheduler based on an SDC formulation that includes a backtracking mechanism to properly handle multiple scheduling constraints and still achieve the minimum possible initiation interval. The SDC formulation has the advantage of being a mathematical framework that supports flexible constraints that are useful for more complex loop pipelines. Furthermore, we describe how to specifically apply associative expression transformations during modulo scheduling to restructure recurrences in complex loops to enable better scheduling. We compared our techniques to existing prior work in modulo scheduling in HLS and also compared against a state-of-art commercial tool. Over a suite of benchmarks, we show that our scheduler and proposed optimizations can result in a geomean wall-clock time reduction of 32% versus prior work and 29% versus a commercial tool. Andrew Canis, Stephen Brown 0003, Jason Helge Anderson |
FPL | 2 |
| 2013 | From software to accelerators with LegUp high-level synthesisabstractEmbedded system designers can achieve energy and performance benefits by using dedicated hardware accelerators. However, implementing custom hardware accelerators for an application can be difficult and time intensive. LegUp is an open-source high-level synthesis framework that simplifies the hardware accelerator design process [8]. With LegUp, a designer can start from an embedded application running on a processor and incrementally migrate portions of the program to hardware accelerators implemented on an FPGA. The final application then executes on an automatically-generated software/hardware coprocessor system. This paper presents on overview of the LegUp design methodology and system architecture, and discusses ongoing work on profiling, hardware/software partitioning, hardware accelerator quality improvements, Pthreads/OpenMP support, visualization tools, and debugging support. Andrew Canis, Jongsok Choi, Blair Fort, Ruolong Lian, Qijing Huang 0001, Nazanin Calagar, Marcel Gort, Jia Jun Qin, Mark Aldham, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
CASES | 11 |
| 2013 | Multi-pumping for resource reduction in FPGA high-level synthesisabstractResource sharing is a classic high-level synthesis (HLS) optimization that saves area by mapping multiple operations to a single functional unit. With resource sharing, only operations scheduled in separate cycles can be assigned to shared hardware, which can result in longer schedules. In this paper, we propose a new approach to resource sharing that allows multiple operations to be performed by a single functional unit in one clock cycle. Our approach is based on multi-pumping, which operates functional units at a higher frequency than the surrounding system logic, typically 2×, allowing multiple computations to complete in a single system cycle. Our approach is particularly effective for DSP blocks on an FPGA, which are used to perform multiply and/or accumulate operations. Our results show that resource sharing using multi-pumping is comparable to traditional resource sharing in terms of area saved, but provides significant performance advantages. Specifically, when targeting a 50% reduction in DSP blocks, traditional resource sharing decreases circuit speed performance by 80%, on average, whereas multi-pumping decreases circuit speed by just 5%. Multi-pumping is a viable approach to achieve the area reductions of resource sharing, with considerably less negative impact to circuit performance. Andrew Canis, Jason Helge Anderson, Stephen Brown 0003 |
DATE | 3 |
| 2013 | The Effect of Compiler Optimizations on High-Level Synthesis for FPGAsabstractWe consider the impact of compiler optimizations on the quality of high-level synthesis (HLS)-generated FPGA hardware. Using a HLS tool implemented within the state-of-the-art LLVM [1] compiler, we study the effect of compiler optimizations on the hardware metrics of circuit area, execution cycles, Fmax, and wall-clock time. We evaluate 56 different compiler optimizations implemented within LLVM and show that some optimizations significantly affect hardware quality. Moreover, we show that hardware quality is also affected by the order in which optimizations are applied. We then present a new HLS-directed approach to compiler optimizations, wherein we execute partial HLS and profiling at intermittent points in the optimization process and use the results to judiciously undo the impact of optimization passes predicted to be damaging to the generated hardware quality. Results show that our approach produces circuits with 16% better speed performance, on average, versus using the standard -O3 optimization level. Qijing Huang 0001, Ruolong Lian, Andrew Canis, Jongsok Choi, Ryan Xi, Stephen Brown 0003, Jason Helge Anderson |
FCCM | 6 |
| 2013 | High-level synthesis with LegUp: a crash course for users and researchersabstractHigh-level synthesis (HLS) has been gaining traction recently as a design methodology for FPGAs, with the promise of raising the productivity of FPGA hardware designers, and ultimately, opening the door to the use of FPGAs as computing devices targetable by software engineers. In this tutorial, we introduce LegUp, an open-source HLS tool for FPGAs developed at the University of Toronto. With LegUp, a user can compile a C program completely to hardware, or alternately, he/she can choose to compile the program to a hybrid hardware/software system comprising a processor along with one or more accelerators. LegUp supports the synthesis of most of the C language to hardware, including loops, structs, multi-dimensional arrays, pointer arithmetic, and floating point operations. The LegUp distribution includes the CHStone HLS benchmark suite, as well as a test suite and associated infrastructure for measuring quality of results, and for verifying the functionality of LegUp-generated circuits. LegUp is freely downloadable at www.legup.org, providing a powerful platform that can be leveraged for new high-level synthesis research. Jason Helge Anderson, Stephen Brown 0003, Andrew Canis, Jongsok Choi |
FPGA | 2 |
| 2013 | From C to Blokus Duo with LegUp high-level synthesisabstractWe apply high-level synthesis (HLS) to generate Blokus Duo game-playing hardware for the FPT 2013 Design Competition [3]. Our design, written in C, is synthesized using the LegUp open-source HLS tool to Verilog, then subsequently mapped using vendor tools to an Altera Cyclone IV FPGA on DE2 board. Our software implementation is designed to be amenable to high-level synthesis, and includes a custom stack implementation, uses only integer arithmetic, and employs the use of bitwise logical operations to improve overall computational performance. The underlying AI decision making is based on alpha-beta pruning [2]. The performance of our synthesizable solution is gauged by playing against the Pentobi [8] - a “known good” C++ software implementation. Jiu Cheng Cai, Ruolong Lian, Andrew Canis, Jongsok Choi, Blair Fort, Eric Hart, Emily Miao, Nazanin Calagar, Stephen Brown 0003, Jason Helge Anderson |
FPT | 11 |
| 2013 | From software threads to parallel hardware in high-level synthesis for FPGAsabstractWe describe the support within high-level hardware synthesis (HLS) for two standard software parallelization paradigms: Pthreads and OpenMP. Parallel code segments, as specified in the software, are automatically synthesized by our HLS tool into parallel-operating hardware sub-circuits. Both data parallelism and task-level parallelism are supported, as is the combined use of both Pthreads and OpenMP. Moreover, our work also provides automated synthesis for commonly occurring synchronization constructs within the Pthreads/OpenMP library: mutual exclusion (mutex) and barriers. Essentially, our framework allows a software engineer to specify parallelism to an HLS tool using methodologies they are likely to be familiar with. An experimental study considers a variety of parallelization scenarios, including demonstrated speedups of up to 12.9× in circuit wall-clock time for the 16-thread case and area-delay product as low as 12% (~8× improvement) when using 4 pipelined hardware threads. Jongsok Choi, Stephen Brown 0003, Jason Helge Anderson |
FPT | 2 |
| 2013 | LegUp: An open-source high-level synthesis tool for FPGA-based processor/accelerator systemsabstractIt is generally accepted that a custom hardware implementation of a set of computations will provide superior speed and energy efficiency relative to a software implementation. However, the cost and difficulty of hardware design is often prohibitive, and consequently, a software approach is used for most applications. In this article, we introduce a new high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. In the hybrid processor/accelerator architecture, program segments that are unsuitable for hardware implementation can execute in software on the processor. LegUp can synthesize most of the C language to hardware, including fixed-sized multidimensional arrays, structs, global variables, and pointer arithmetic. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. We also give results demonstrating the ability of the tool to explore the hardware/software codesign space by varying the amount of a program that runs in software versus hardware. LegUp, along with a set of benchmark C programs, is open source and freely downloadable, providing a powerful platform that can be leveraged for new research on a wide range of high-level synthesis topics. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2013 | Exploiting Task- and Data-Level Parallelism in Streaming Applications Implemented in FPGAsabstractThis article describes the design and implementation of a novel compilation flow that implements circuits in FPGAs from a streaming programming language. The streaming language supported is called FPGA Brook and is based on the existing Brook language. It allows system designers to express applications in a way that exposes parallelism, which can be exploited through hardware implementation. FPGA Brook supports replication, allowing parts of an application to be implemented as multiple hardware units operating in parallel. Hardware units are interconnected through FIFO buffers which use the small memory modules available in FPGAs. The FPGA Brook automated design flow uses a source-to-source compiler, developed as a part of this work, and combines it with a commercial behavioral synthesis tool to generate the hardware implementation. A suite of benchmark applications was developed in FPGA Brook and implemented using our design flow. Experimental results indicate that performance of many applications scales well with replication. Our benchmark applications also achieve significantly better results than corresponding implementations using a commercial behavioral synthesis tool. We conclude that using an automated design flow for implementation of streaming applications in FPGAs is a promising methodology. Franjo Plavec, Zvonko G. Vranesic, Stephen Brown 0003 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2012 | Impact of Cache Architecture and Interface on Performance and Area of FPGA-Based Processor/Parallel-Accelerator SystemsabstractWe describe new multi-ported cache designs suitable for use in FPGA-based processor/parallel-accelerator systems, and evaluate their impact on application performance and area. The baseline system comprises a MIPS soft processor and custom hardware accelerators with a shared memory architecture: on-FPGA L1 cache backed by off-chip DDR2 SDRAM. Within this general system model, we evaluate traditional cache design parameters (cache size, line size, associativity). In the parallel accelerator context, we examine the impact of the cache design and its interface. Specifically, we look at how the number of cache ports affects performance when multiple hardware accelerators operate (and access memory) in parallel, and evaluate two different hardware implementations of multi-ported caches using: 1) multi-pumping, and 2) a recently-published approach based on the concept of a live-value table. Results show that application performance depends strongly on the cache interface and architecture: for a system with 6 accelerators, depending on the cache design, speed up swings from 0.73× to 6.14×, on average, relative to a baseline sequential system (with a single accelerator and a direct-mapped, 2KB cache with 32B lines). Considering both performance and area, the best architecture is found to be a 4-port multi-pump direct-mapped cache with a 16KB cache size and a 128B line size. Jongsok Choi, Kevin Nam, Andrew Canis, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FCCM | 5 |
| 2012 | Impact of FPGA architecture on resource sharing in high-level synthesisabstractResource sharing is a key area-reduction approach in high-level synthesis (HLS) in which a single hardware functional unit is used to implement multiple operations in the high-level circuit specification. We show that the utility of sharing depends on the underlying FPGA logic element architecture and that different sharing trade-offs exist when 4-LUTs vs. 6-LUTs are used. We further show that certain multi-operator patterns occur multiple times in programs, creating additional opportunities for sharing larger composite functional units comprised of patterns of interconnected operators. A sharing cost/benefit analysis is used to inform decisions made in the binding phase of an HLS tool, whose RTL output is targeted to Altera commercial FPGA families: Stratix IV (dual-output 6-LUTs) and Cyclone II (4-LUTs). Stefan Hadjis, Andrew Canis, Jason Helge Anderson, Jongsok Choi, Kevin Nam, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 6 |
| 2011 | Low-cost hardware profiling of run-time and energy in FPGA embedded processorsabstractThis paper introduces a low-overhead hardware profiling architecture, called LEAP, that attains real-time cycle and energy profiles of an FPGA-based soft processor. A novel technique is used to associate profiling data with specific functions in a way that is area- and power-efficient. Results show that relative to a previously-published hardware profiler, our design uses up to 18× less area and 8.6× less energy. LEAP is designed to be extensible for a variety of profiling tasks, three of which are investigated in this paper. We also demonstrate the utility of LEAP in the context of hardware/software co-design of processor/accelerator FPGA-based systems. Mark Aldham, Jason Helge Anderson, Stephen Brown 0003, Andrew Canis |
ASAP | 3 |
| 2011 | LegUp: high-level synthesis for FPGA-based processor/accelerator systemsabstractIn this paper, we introduce a new open source high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 7 |
| 2011 | Toward Automated ECOs in FPGAsabstractEngineering change orders (ECOs), which are used to apply late-stage specification changes and bug fixes, have become an important part of the field-programmable gate array design flow. ECOs are beneficial since they are applied directly to a placed-and-routed netlist which preserves most of the engineering effort invested previously. Unfortunately, designers often apply ECOs in a manual fashion which may have an unpredictable impact on the design's final correctness and end costs. As a solution, this paper introduces an automated method to tackle the ECO problem. This paper uses a novel resynthesis technique which can automatically update the functionality of a circuit by leveraging the existing logic within the design, thereby removing the inefficient manual effort required by a designer. The technique presented in this paper is robust enough to handle a wide range of changes. Furthermore, the technique can successfully make late-stage functional changes while minimally perturbing the placed-and-routed netlist: something that is necessary for ECOs. Also, this technique does this with a minimal impact on the circuit performance where on average over 90% of the placement and routing wires remain unchanged. Andrew C. Ling, Stephen Brown 0003, Sean Safarpour, Jianwen Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Technology issues facing the world's largest integrated circuitsabstractSummary form only given. FPGAs are amongst the world's largest and most complex integrated circuits, and they continue to be very early adopters of the latest process technology. This talk will describe some of the driving applications and technology trends pushing FPGAs to 28 nm and smaller process nodes. We will also highlight how FPGA architecture is evolving, as exemplified by Altera's Stratix V FPGAs. Power management and silicon efficiency issues are pushing FPGAs to become somewhat more application-targeted, and to incorporate larger amounts of hard logic that makes them more complete systems-on-achip. In addition, the very high I/O bandwidth requirements of next-generation systems are driving innovation in both high-speed memory interface design and high-speed serial transceiver design. Stratix V supports partial reconfiguration to increase silicon efficiency by swapping in different functionality over time. We will describe both the hardware that enables partial reconfiguration, and the software tools that will enable efficient design without becoming entangled in low-level physical details. Finally, we will discuss both software challenges and promising research efforts to create CAD tools that will help designers productively create the very large systems enabled by modern FPGAs. Stephen Brown 0003 |
FPT | 1 |
| 2010 | Decomposition-Based Vectorless Toggle Rate Computation for FPGA CircuitsabstractThis paper presents a novel and accurate method of estimating the toggle rates of signals in field-programmable gate array (FPGA)-based logic circuits without the use of simulation vectors. Compared to previous vectorless techniques, our approach provides improved accuracy-of-results, especially for individual signals, which could be leveraged by computer-aided design (CAD) tools for performing power optimization of logic circuits. Increased accuracy is achieved by using stochastic methods that estimate the transition densities at FPGA logic elements while accounting for both spatial and temporal correlation of logic signals. Spatial correlation is calculated by leveraging a unique XOR-based decomposition technique that provides both accurate results and fast computation times. We also consider the delay information of implemented circuits, providing for a comprehensive treatment of glitches, including the effects of inertial limits on power dissipation. Our toggle-rate estimation approach has been tested on a commonly used set of Microelectronic Center of North Carolina circuits, as well as a set of industrial circuits targeted to Altera Stratix II FPGAs. Results show that our techniques provide a three times lower percent error, while maintaining a low processing time, when compared to two existing techniques: the vectorless estimation tool shipped with the commercial Quartus II 8.0 CAD tool, and the ACE v2.0 academic tool produced from the University of British Columbia, Vancouver, BC, Canada. Tomasz S. Czajkowski, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Towards automated ECOs in FPGAsabstractDuring the FPGA design flow, engineering change orders (ECOs) have become an essential methodology to apply late-stage specification changes and bug fixes. ECOs are beneficial since they are applied directly to a place-and-routed netlist which preserves most of the engineering effort invested previously. Unfortunately, designers often apply ECOs in a manual fashion which has an unpredictable impact on the design's final correctness and end costs. As a solution, we introduce an automated method to tackle the ECO problem. Specifically, we introduce a resynthesis technique which can automatically update the functionality of a circuit by leveraging the existing logic within the design; thereby removing the inefficient manual effort required by a designer. Our technique is robust enough to handle a wide range of changes. Furthermore, our technique can successfully make late-stage functional changes while minimally perturbing the place-and-routed netlist: something that is necessary for ECOs. When applied to several benchmarks on Altera's Stratix architecture, we show that our approach can automatically apply ECOs in over 80% of the cases presented. Furthermore, our technique does this with a minimal impact to the circuit performance where on average over 90% of the placement and routing wires remain unchanged. Andrew C. Ling, Stephen Brown 0003, Jianwen Zhu, Sean Safarpour |
FPGA | 2 |
| 2009 | Enhancements to FPGA design methodology using streamingabstractCapacity of FPGAs has grown significantly, leading to increased complexity of designs targeting these chips. Traditional FPGA design methodology using HDLs is no longer sufficient and new methodologies are being sought. An attractive possibility is to use streaming languages. Streaming languages group data into streams, which are processed by computational nodes called kernels. They are suitable for implementation in FPGAs because they expose parallelism, which can be exploited by implementing the application in FPGA logic. Designers can express their designs in a streaming language and target FPGAs without needing a detailed understanding of digital logic design. In this paper we show how the Brook streaming language can be used to simplify design for FPGAs, while providing reasonable performance compared to other methodologies. We show that throughput of streaming applications can be increased through automatic kernel replication. Using our compiler, the FPGA designer can trade off FPGA area and performance by changing the amount of kernel replication. We describe the details of our compiler and present performance and area of a set of benchmarks. We found that throughput scales well with increased replication for most applications. Franjo Plavec, Zvonko G. Vranesic, Stephen Brown 0003 |
FPL | 3 |
| 2008 | Functionally linear decomposition and synthesis of logic circuits for FPGAsabstractThis paper presents a novel logic synthesis method to reduce the area of XOR-based logic functions. The idea behind the synthesis method is to exploit linear dependency between logic sub-functions to create an implementation based on an XOR relationship with a lower area overhead. Experiments conducted on a set of 99 MCNC benchmark (25 XOR based, 74 non-XOR) circuits show that this approach provides an average of 18.8% area reduction as compared to BDS-PGA 2.0 and 25% area reduction as compared to ABC for XOR-based logic circuits. Tomasz S. Czajkowski, Stephen Brown 0003 |
DAC | 2 |
| 2008 | Towards Compilation of Streaming Programs into FPGA HardwareabstractThere is an increasing need for automated conversion of high-level design descriptions into hardware. We present a flow that converts a software application written in the Brook streaming language into a hardware description targeting FPGAs. We use a combination of our source-to-source compiler and a commercial C2H behavioral synthesis compiler. Our approach results in a significant through-put increase compared to software and ordinary C2H results (up to 8.9X and 4.3X, respectively). The throughput can be further increased by using more hardware resources to exploit data parallelism available in streaming applications. Franjo Plavec, Zvonko G. Vranesic, Stephen Brown 0003 |
FDL | 3 |
| 2008 | Fast toggle rate computation for FPGA circuitsabstractThis paper presents a fast and scalable method of computing signal toggle rate in FPGA-based circuits. Our technique is a vectorless estimation technique, which can be used in a CAD tool to identify the parts of the circuit that can benefit from power optimization. A key advantage of our approach is its ability to efficiently account for spatial correlation of related logic cones, which is accomplished using a novel XOR-based decomposition. In addition, our approach uses post-routing circuit delays to account for glitches in a logic circuit. The proposed approach was tested on 14 MCNC benchmark circuits compiled for the Altera Stratix II devices. The results indicate that our method improves the vectorless estimation technique available in the latest version of Alterapsilas Quartus II commercial CAD tool, reducing the average error by 37% and standard deviation by 59%. Tomasz S. Czajkowski, Stephen Brown 0003 |
FPL | 2 |
| 2008 | Delay driven AIG restructuring using slack budget managementabstractTiming optimizations during logic synthesis has become a necessary step to achieve timing closure in VLSI designs. This often involves "shortening" all paths found in the circuit at a cost of increasing the circuit area. In contrast, we present a synthesis approach which leverages slack budgeting to effectively minimize the critical path length without increasing the area of the design. Our results confirm that this is an effective method to control area while optimizing for delay. When compared to an area driven logic synthesis flow, we achieve a 32% reduction in logic depth and an 11% reduction in circuit delay when placed by VPR [1]; and when compared against a depth controlled logic synthesis flow without slack budgeting, we achieve an 8% reduction in logic depth and a 3% reduction in circuit delay when placed by VPR [1]. In both cases, the area penalty is negligible. Andrew C. Ling, Jianwen Zhu, Stephen Brown 0003 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Functionally Linear Decomposition and Synthesis of Logic Circuits for FPGAsabstractThis paper presents a novel XOR-based logic synthesis approach called functionally linear decomposition and synthesis (FLDS). This approach decomposes a logic function to expose an XOR relationship by using Gaussian elimination. It is fundamentally different from the traditional approaches to this problem, which are based on the work of Ashenhurst and Curtis. FLDS utilizes binary decision diagrams to efficiently represent logic functions, making it fast and scalable. This technique was tested on a set of 99 MCNC benchmarks, mapping each design into a network of four input lookup tables. On the 25 of the benchmarks, which have been classified by previous researchers as XOR-based logic circuits, our approach provides significant area savings. In comparison to the leading logic synthesis tools, ABC and BDS-PGA 2.0, FLDS produces XOR-based circuits with 25.3% and 18.8% smaller area, respectively. The logic circuit depth is also improved by 7.7% and 14.5%, respectively. Tomasz S. Czajkowski, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Scalable Synthesis and Clustering Techniques Using Decision DiagramsabstractBinary-decision diagrams (BDDs) have proven to be an efficient means to represent and manipulate Boolean formulas and sets due to their compactness and canonicity. In this paper, we leverage the efficiency of BDDs for new areas in field-programmable gate-array (FPGA) computer-aided design (CAD) flow including cut generation and clustering by reducing these problems to BDDs and solving them using Boolean operations. As a result, we show that this leads to more than 10 reduction in runtime and memory use when compared to previous techniques, as reported by Mishchenko and Lin. This speedup allows us to apply this paper to new areas in the FPGA CAD flow previously not possible. Specifically, we introduce a new method to solve the logic-synthesis elimination problem found in FBDD, a reported BDD synthesis engine with an order-of-magnitude speedup over SIS. Our new elimination algorithm results in an overall speedup of 6 in FBDD with no impact on circuit area. Andrew C. Ling, Jianwen Zhu, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | BddCut: Towards Scalable Symbolic Cut EnumerationabstractWhile the covering algorithm has been perfected recently by the iterative approaches, such as DAOmap and IMap, its application has been limited to technology mapping. The main factor preventing the covering problem's migration to other logic transformations, such as elimination and resynthesis region identification found in SIS and FBDD, is the exponential number of alternative cuts that have to be evaluated. Traditional methods of cut generation do not scale beyond a cut size of 6. In this paper, a symbolic method that can enumerate all cuts is proposed without any pruning, up to a cut size of 10. We show that it can outperform traditional methods by an order of magnitude and, as a result, scales to 100K gate benchmarks. As a practical driver, the covering problem applied to elimination is shown where it can not only produce competitive area, but also provide more than 6times average runtime reduction of the total runtime in FBDD, a BDD based logic synthesis tool with a reported order of magnitude faster runtime than SIS and commercial tools with negligible impact on area. Andrew C. Ling, Jianwen Zhu, Stephen Brown 0003 |
ASP-DAC | 3 |
| 2007 | Using Negative Edge Triggered FFs to Reduce Glitching Power in FPGA CircuitsabstractThis paper presents an algorithm for reducing dynamic power dissipated by Field-Programmable Gate Array (FPGA) circuits. The algorithm uses a fast probability based model to estimate glitches on wires in a circuit and then inserts negative edge triggered FFs at outputs of Lookup Tables (LUTs) that produce glitches. A negative edge triggered FF maintains the logic value produced by the LUT in the previous cycle for the first half of the clock period, filtering glitches that occur at the output of the LUT. The power dissipation is lowered by reducing the number of transitions that propagate to the general routing network. Tomasz S. Czajkowski, Stephen Brown 0003 |
DAC | 2 |
| 2007 | Incremental placement for structured ASICs using the transportation problemabstractWhile physically driven synthesis techniques have proven to be an effective method to meet tight timing constraints required by a design, the incremental placement step during physically driven synthesis has emerged as the primary bottleneck. As a solution, this paper introduces a scalable incremental placement algorithm based upon the well known transportation problem. This method has an average speedup of 2× and a 30% reduction in memory usage when compared against a commercial incremental placer without any impact on area or speed of the final placed circuit. Furthermore, this method is scalable for structured ASICs. Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003 |
VLSI-SoC | 3 |
| 2007 | An area-efficient timing closure technique for FPGAs using Shannon's expansion
Deshanand P. Singh, Stephen Brown 0003 |
Integr. | 2 |
| 2007 | FPGA PLB Architecture Evaluation and Area Optimization Techniques Using Boolean SatisfiabilityabstractThis paper presents a field-programmable gate array (FPGA) logic synthesis technique based upon Boolean satisfiability. This paper shows how to map any Boolean function into an arbitrary programmable logic block (PLB) architecture without any custom decomposition techniques. The authors illustrate several useful applications of this technique by showing how this technique can be used for architecture evaluation and area optimization. When evaluating the FPGA architecture, the authors focus on the basic building block of the FPGA, which they refer to as PLB. In order to illustrate the flexibility of their evaluation framework, several unrelated PLB architectures are evaluated in an automated fashion. Furthermore, the authors show that using their technique is able to reduce FPGA resource usage by 27% on average in common subcircuits found in digital design. Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | Predicting Interconnect Delay for Physical Synthesis in a FPGA CAD FlowabstractThis paper studies the prediction of interconnect delay in an industrial setting. Industrial circuits and two industrial field-programmable gate-array (FPGA) architectures were used in this paper. We show that there is a large amount of inherent randomness in a state-of-the-art FPGA placement algorithm. Thus, it is impossible to predict interconnect delay with a high degree of accuracy. Furthermore, we show that a simple timing model can be used to predict some aspects of interconnect timing with just as much accuracy as predictions obtained by running the placement tool itself. Using this simple timing model in a two-phase timing driven physical synthesis flow can both improve quality of results and decrease runtime. Next, we present a metric for predicting the accuracy of our interconnect delay model and show how this metric can be used to reduce the runtime of a timing driven physical synthesis flow. Finally, we examine the benefits of using the simple timing model in a timing driven physical synthesis flow, and attempt to establish an upper bound on these possible gains, given the difficulty of interconnect delay prediction. Valavan Manohararajah, Gordon R. Chiu, Deshanand P. Singh, Stephen Brown 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | A Multithreaded Soft Processor for SoPC Area ReductionabstractThe growth in size and performance of field programmable gate arrays (FPGAs) has compelled system-on-a-programmable-chip (SoPC) designers to use soft processors for controlling systems with large numbers of intellectual property (IP) blocks. Soft processors control IP blocks, which are accessed by the processor either as peripheral devices or/and by using custom instructions (CIs). In large systems, chip multiprocessors (CMPs) are used to execute many programs concurrently. When these programs require the use of the same IP blocks which are accessed as peripheral devices, they may have to stall waiting for their turn. In the case of CIs, the FPGA logic blocks that implement the CIs may have to be replicated for each processor. In both of these cases FPGA area is wasted, either by idle soft processors or the replication of CI logic blocks. This paper presents a multithreaded (MT) soft processor for area reduction in SoPC implementations. An MT processor allows multiple programs to access the same IP without the need for the logic replication or the replication of whole processors. We first designed a single-threaded processor that is instruction-set compatible to Altera's Nios II soft processor. Our processor is approximately the same size as the Nios II economy version, with equivalent performance. We augmented our processor to have 4-way interleaved multithreading capabilities. This paper compares the area usage and performance of the MT processor versus two CMP systems, using Altera's and our single-threaded processors, separately. Our results show that we can achieve an area savings of about 45% for the processor itself, in addition to the area savings due to not replicating CI logic blocks Blair Fort, Davor Capalija, Zvonko G. Vranesic, Stephen Brown 0003 |
FCCM | 4 |
| 2006 | Modular Partitioning for Incremental CompilationabstractThis paper presents an automated partitioning strategy to divide a design into a set of partitions based on design hierarchy information. While the primary objective is to use these partitions in an incremental design flow for compile time reduction, the performance of the partitioned design should not be degraded after partitioning. Experimental results using the incremental design feature of Altera's Quartus tool show that our algorithm can generate partitioning solutions comparable with a set of manually partitioned industrial circuits and results in more than 50% compile time reduction Mehrdad Eslami Dehkordi, Stephen Brown 0003, Terry P. Borer |
FPL | 2 |
| 2006 | Adaptive FPGAs: High-Level Architecture and a Synthesis MethodabstractThis paper presents preliminary work exploring adaptive field programmable gate arrays (AFPGAs). An AFPGA is adaptive in the sense that the functionality of subcircuits placed on the chip can change in response to changes observed on certain control signals. We describe the high-level architecture which adds additional control logic and SRAM bits to a traditional FPGA to produce an AFPGA. We also describe a synthesis method that identifies and resynthesizes mutually exclusive pieces of logic so that they may share the resources available in an AFPGA. The architectural feature and its associated synthesis method helps reduce circuit size by 28% on average and up to 40% on select circuits Valavan Manohararajah, Stephen Brown 0003, Zvonko G. Vranesic |
FPL | 2 |
| 2006 | Mapping arbitrary logic functions into synchronous embedded memories for area reduction on FPGAsabstractThis work describes a new mapping technique, RAM-MAP, that identifies parts of circuits that can be efficiently mapped into the synchronous embedded memories found on field programmable gate arrays (FPGAs). Previous techniques developed for mapping into asynchronous embedded memories cannot be used because modern FPGAs do not have asynchronous embedded memories. After technology mapping, an area-prediction cost function is used to guide the selection of logic cones to be placed in embedded memories. Extra logic is added to compensate for missing asynchronous functionality on the synchronous memories. Experiments conducted on Altera's Stratix device family indicate that this embedded memory mapping technique can provide an average area reduction of 6.2% and up to 32.5% on a large set of industrial designs. A small architecture change that increases the size of the FPGA fabric by 0.05% can increase the average area reduction to 14.1% and up to 59.1% on the same design set. Gordon R. Chiu, Deshanand P. Singh, Valavan Manohararajah, Stephen Brown 0003 |
ICCAD | 4 |
| 2006 | Heuristics for Area Minimization in LUT-Based FPGA Technology MappingabstractIn this paper, an iterative technology-mapping tool called IMap is presented. It supports depth-oriented (area is a secondary objective), area-oriented (depth is a secondary objective), and duplication-free mapping modes. The edge-delay model (as opposed to the more commonly used unit-delay model) is used throughout. Two new heuristics are used to obtain area reductions over previously published methods. The first heuristic predicts the effects of various mapping decisions on the area of the final solution, and the second heuristic bounds the depth of the mapping solution at each node. In depth-oriented mode, when targeting five lookup tables (LUTs), IMap obtains depth optimal solutions that are 44.4%, 19.4%, and 5% smaller than those produced by FlowMap, CutMap, and DAOMap, respectively. Targeting the same LUT size in area-oriented mode, IMap obtains solutions that are 17.5% and 9.4% smaller than those produced by duplication-free mapping and ZMap, respectively. IMap is also shown to be highly efficient. Runtime improvements of between 2.3times and 82times are obtained over existing algorithms when targeting five LUTs. Area and runtime results comparing IMap to the other mappers when targeting four and six LUTs are also presented Valavan Manohararajah, Stephen Brown 0003, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | FPGA technology mapping: a study of optimalityabstractThis paper attempts to quantify the optimality of FPGA technology mapping algorithms. We develop an algorithm, based on Boolean satisfiability (SAT), that is able to map a small subcircuit into the smallest possible number of lookup tables (LUTs) needed to realize its functionality. We iteratively apply this technique to small portions of circuits that have already been technology mapped by the best available mapping algorithms for FPGAs. In many cases, the optimal mapping of the subcircuit uses fewer LUTs than is obtained by the technology mapping algorithm. We show that for some circuits the total area improvement can be up to 67%. Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003 |
DAC | 3 |
| 2005 | Incremental retiming for FPGA physical synthesisabstractIn this paper, we present a new linear-time retiming algorithm that produces near-optimal results. Our implementation is specically targeted at Altera's Stratix [1] FPGA-based designs, although the techniques described are general enough for any implementation medium. The algorithm is able to handle the architectural constraints of the target device, multiple timing constraints assigned by the user and implicit legality constraints. It ensures that register moves do not create asynchonous problems such as creating a glitch on a clock/reset signal. Deshanand P. Singh, Valavan Manohararajah, Stephen Brown 0003 |
DAC | 3 |
| 2005 | FPGA PLB Evaluation using Quantified Boolean SatisfiabilityabstractThis paper describes a novel field programmable gate array (FPGA) logic synthesis technique which determines if a logic function can be implemented in a given programmable circuit and describes how this problem can be formalized and solved using quantified Boolean satisfiability. This technique is general enough to be applied to any type of logic function and programmable circuit; thus, it has many applications to FPGAs. The application demonstrated in this paper is FPGA PLB evaluation where their results show that this tool allows radical new features of FPGA logic blocks to be evaluated in a rigorous scientific way. Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003 |
FPL | 3 |
| 2005 | Post-Placement BDD-Based Decomposition for FPGAsabstractThis work explores the effect of adding a timing driven functional decomposition step to the traditional field programmable gate array (FPGA) CAD flow. Once placement has completed, alternative decompositions of the logic on the critical path are examined for potential delay improvements. The placed circuit is then modified to use the best decompositions found. Any placement illegalities introduced by the new decompositions are resolved by an incremental placement step. Experiments conducted on Altera's Stratix and Stratix II device families indicate that this functional decomposition technique can provide average performance improvements of 6.1% and 5.6% on a large set of industrial designs, respectively. Valavan Manohararajah, Deshanand P. Singh, Stephen Brown 0003 |
FPL | 3 |
| 2005 | FPGA Logic Synthesis Using Quantified Boolean Satisfiability
Andrew C. Ling, Deshanand P. Singh, Stephen Brown 0003 |
SAT | 3 |
| 2004 | Retiming aware clustering for sequential circuitsabstractThis work presents a simultaneous sequential retiming and clustering algorithm for delay minimization applicable to FPGAs. The algorithm is based on Pan et al.(1998) with several modifications and enhancements to improve the performance of the final clustered circuits. A duplication control strategy is used to reduce the amount of node duplication. Experimental results on the biggest MCNC benchmark circuits using Altera's Quartus show that our algorithm can increase the performance, on average, by almost 22% compared with the case when Quartus is used without our clustering information. Mehrdad Eslami Dehkordi, Stephen Brown 0003 |
FPT | 2 |
| 2004 | The Quartus University Interface Program: enabling advanced FPGA researchabstractFPGA researchers constantly face the challenge of determining whether their innovations will work in the real world. The Quartus University Interface Program (QUIP) allows the researcher to answer this important question by directly integrating research prototypes within one of the FPGA industry's leading CAD tool suites. This work describes the QUIP interface as well as research projects that are of significant interest to the FPGA industry. Shawn Malhotra, Terry P. Borer, Deshanand P. Singh, Stephen Brown 0003 |
FPT | 4 |
| 2003 | Recursive circuit clustering for minimum delay and areaabstractWe present an effective recursive algorithm for circuit clustering for delay and area minimization, which is applicable to FPGAs. At the highest level of clustering, the circuit is clustered using a modified single-level clustering algorithm. A cluster to netlist transformation technique is proposed, which converts each cluster into a new subcircuit. The algorithm then continues recursively by clustering the generated subcircuits into further levels of clusters. To reduce the amount of node duplication and the number of clusters at each level of clustering, we propose a node removal algorithm based on the node slack along with a simple cluster-packing algorithm. Experimental results on the two-level clustering problem using Quartus Design System from Altera show that our algorithm achieves, on average, 7.3% more delay reduction when compared to the latest published work on the problem. Also total FPGA compile time reported by Quartus is reduced by 36%. Mehrdad Eslami Dehkordi, Stephen Brown 0003 |
FPGA | 2 |
| 2003 | Performance-driven recursive multi-level clusteringabstractThis paper presents an effective algorithm for multi-level circuit clustering for delay minimization, which is applicable to FPGAs. At the highest level of clustering, the circuit is clustered using a modified single-level clustering algorithm. A cluster to netlist transformation technique is proposed, which converts each cluster into a new subcircuit. The algorithm then continues recursively by clustering the generated sub-circuits into further levels of clusters. To reduce the amount to area overhead, a node duplication control algorithm based on the node slack is proposed and a cluster packing algorithm is used to reduce the number of clusters. Experimental results on the two-level clustering problem using Quartus Design System from Altera show that our algorithm reduces the delay, on average, by 7.3% compared with the best results in. Also the total FPGA compile time for all the benchmark circuits reported by Quartus is reduced by 36%. Mehrdad Eslami Dehkordi, Stephen Brown 0003 |
FPT | 2 |
| 2002 | Integrated retiming and placement for field programmable gate arraysabstractRetiming is a synchronous circuit transformation that can optimize the delay of a synchronous circuit by moving registers across combinational circuit elements. The combinational structure remains unchanged and the observable behavior of the circuit is identical to the original.In this paper, we address the problem of applying retiming techniques to circuits implemented in Field Programmable Gate Arrays (FPGAs). FPGAs contain prefabricated and configurable routing elements that allow us to easily implement a variety of circuits. However this interconnect contributes greatly to the overall delay in the implemented circuit. If a circuit is retimed prior to the placement and routing phases of the CAD flow, then it has no information about the delays introduced by the configurable interconnect. Our fundamental experiment is to determine whether there are any gains in tightly coupling retiming and placement so that the retiming algorithm has some estimate of the routing delays.Specifically, we introduce a post-placement retiming algorithm that understands how to take advantage of FPGA architectural features. This retiming algorithm may introduce extra registers into the circuit. These new registers need to be placed in some location in the FPGA. Retiming register placement is accomplished by a novel incremental clustering and placement algorithm. The incremental algorithm builds upon the placement of the non-retimed circuit to intelligently sift in the newly-introduced registers.In addition, we explore making the placement algorithms "retiming aware." These placement algorithms try to place logic blocks in such a way that the subsequent retiming produces better speed results. These techniques include the identification of retiming-critical cycles during placement.Our experiments show that the integration of retiming with placement results in 19% better clock periods in comparison to the application of retiming before the place and route steps. Deshanand P. Singh, Stephen Brown 0003 |
FPGA | 2 |
| 2002 | Constrained clock shifting for field programmable gate arraysabstractCircuits implemented in FPGAs have delays that are dominated by its programmable interconnect. This interconnect provides the ability to implement arbitrary connections. However, it contains both highly capacitive and resistive elements. The delay encountered by any connection depends strongly on the number of interconnect elements used to route the connection. These delays are only completely known after the place and route phase of the CAD flow. We propose the use of Clock Shifting optimization techniques to improve the clock frequency as a post place and route step.Clock Shifting Optimization is a technique first formalized in [4]. It is a cycle-stealing algorithm that allows one to reduce the critical path delay of a synchronous circuit by shifting the clock signals at each register. This technique allows late arriving signals to be sampled at a later point in time by intentionally introducing a skew on the clock input of the sampling register. Typical FPGAs contain a number of special purpose global clock networks that distribute clock signals to every register in the chip. Unused global clock lines in FPGAs can be used to distribute a finite set of clock skews to the entire circuit. We propose an efficient integer programming method to find the optimal circuit improvement for a finite set of clock skews. This technique is modified to consider inherent uncertainties present in the timing models. The uncertainty controls the aggressiveness of the optimizations as we must take great care in ensuring functionality for any range of possible timing characteristics.Our results confirm intuition that more aggressive speed optimizations can be performed as timing models become more accurate. We also show that providing 4 skewed versions of the nominal clock signal results in the best delay--area tradeoff. This result is evocative as it may suggest future FPGA architectures that contain greater numbers of global clock lines, as we tradeoff gains in speed for greater power requirements from increased clock network flexibility. Deshanand P. Singh, Stephen Brown 0003 |
FPGA | 2 |
| 2002 | Automatic Partitioning for Improved Placement and Routing in Complex Programmable Logic Devices
Valavan Manohararajah, Terry P. Borer, Stephen Brown 0003, Zvonko G. Vranesic |
FPL | 3 |
| 2002 | The effect of cluster packing and node duplication control in delay driven clusteringabstractAlthough delay driven clustering algorithms can optimize circuit delay, they usually result in huge area increase. We present a node duplication control strategy along with a simple packing algorithm that greatly reduce the area penalty with a very small degradation in performance. We use the Quartus Design System from Altera to test our algorithm for a set of MCNC benchmark circuits. The results show that while a 14.5% average delay decrease can be achieved with an average area increase of 240.8%, the algorithm reduces the average area penalty to 27.4% with an average delay decrease of 13.8%. Also the number of clusters and the fitting time reported by Quartus are reduced by more than 91% and 60%, respectively. Mehrdad Eslami Dehkordi, Stephen Brown 0003 |
FPT | 2 |
| 2002 | Incremental placement for layout driven optimizations on FPGAsabstractThis paper presents an algorithm to update the placement of logic elements when given an incremental netlist change. Specifically, these algorithms are targeted to incrementally place logic elements created by layout-driven circuit restructuring techniques. The incremental placement engine assumes that the restructuring algorithms provide a list of new logic elements along with preferred locations for each of these new elements. It then tries to shift non-critical logic elements in the original placement out of the way to satisfy the preferred location requests. Our algorithm considers modern FPGA architectures with clustered logic blocksthat have numerous architectural constraints. Experiments indicate that our technique produces results of extremely highquality. Deshanand P. Singh, Stephen Brown 0003 |
ICCAD | 2 |
| 2001 | The case for registered routing switches in field programmable gate arraysabstractFPGAs are characterized by a programmable interconnect that contains highly resistive and capacitive elements. While the configurable structure of the interconnect allows for the implementation of arbitrary circuits, it has also become a significant bottleneck for high-speed circuits. Even if there are only a few signal paths that run along long stretches of interconnect, it is these paths that may determine the maximum operating frequency of the circuit. Deshanand P. Singh, Stephen Brown 0003 |
FPGA | 2 |
| 2000 | Technology mapping issues for an FPGA with lookup tables and PLA-like blocksabstractIn this paper we present new technology mapping algorithms for use in a programmable logic device (PLD) that contains both lookup tables (LUTs) and PLA-like blocks. The technology mapping algorithms partially collapse circuits to reduce either area or depth, and pack the circuits into a minimum number of LUTs and PLA-like blocks. Since no other technology mapping algorithm for this problem has been previously published, we cannot compare our approach to others. Instead, to illustrate the importance of this problem we use our algorithms to investigate the benefits provided by a PLD architecture with both LUTs and PLA-like blocks compared to a traditional LUT-based FPGA. The experimental results indicate that our mixed PLD architecture is more area-efficient than LUT-based FPGAs by up to 29%, or more depth-efficient by up to 75%.1 Alireza Kaviani, Stephen Brown 0003 |
FPGA | 2 |
| 2000 | The NUMAchine MultiprocessorabstractSmall-scale multiprocessors are becoming increasingly economical and common, whereas larger multiprocessors continue to have higher per-node costs. The NUMAchine multiprocessor project seeks to make large-scale multiprocessors more economical while maintaining high performance by exploring architectural and hardware features for low-cost, modular multiprocessors. To demonstrate our approach, we have implemented a prototype system that is scalable to 128 processors. An efficient directory-based cache coherence protocol exploits our hierarchical ring-based interconnect and supports sequential consistency. This paper documents the design choices and the resulting performance of the system using both simulation results and measurements on the prototype hardware. R. Grindley, Tarek S. Abdelrahman, Stephen Brown 0003, S. Caranci, D. DeVries, Benjamin Gamsa, A. Grbic, M. Gusat, R. Ho, Orran Krieger, Guy Lemieux, K. Loveless, Naraig Manjikian, P. McHardy, Sinisa Srbljic, Michael Stumm, Zvonko G. Vranesic, Zeljko Zilic |
ICPP | 3 |
| 1998 | Technology Mapping for Large Complex PLDsabstractIn this paper we present a new technology mapping algorithm for use with complex PLDs (CPLDs), which consists of a large number of PLA-style logic blocks. Although the traditional synthesis approach for such devices uses two-level minimization, the complexity of recently-produced CPLDs has resulted in a trend toward multi-level synthesis. We describe an approach that allows existing multi-level synthesis techniques [13] to be adapted to produce circuits that are well-suited for implementation in CPLDs. Our algorithm produces circuits that require up to 90% fewer logic blocks than the circuits produced by a recently-published algorithm. Jason Helge Anderson, Stephen Brown 0003 |
DAC | 2 |
| 1998 | Design and Implementation of the NUMAchine MultiprocessorabstractThis paper describes the design and implementation of the NUMAchine multiprocessor. As the market for CC-NUMA multiprocessors expands, this research project provides a timely architectural design and cost-effective prototype. The key to the successful implementation of our 48-processor prototype is the use of off-the-shelf components and programmable logic devices. Since this machine will serve as a research vehicle for parallel software development, a number of hardware features to enhance experimentation have been included in the design. A. Grbic, Stephen Brown 0003, S. Caranci, R. Grindley, M. Gusat, Guy Lemieux, K. Loveless, Naraig Manjikian, Sinisa Srbljic, Michael Stumm, Zvonko G. Vranesic, Zeljko Zilic |
DAC | 2 |
| 1998 | An LPGA with Foldable PLA-style Logic BlocksabstractLaser-programmed gate arrays (LPGAs) represent a new approach to application specific integrated circuit prototyping and implementation. This paper proposes a new LPGA logic block architecture called a foldable PLA-style logic block. The proposed logic block architecture is similar to that found in commercially available CPLDs. The term foldable means that the granularity of the logic block can be varied. This is achieved using the LPGA laser disconnect methodology. A custom CAD tool has been developed to map circuits into the new logic block architecture. An experimental study shows that LPGAs with foldable logic blocks are more area-efficient than those based on normal unfoldable logic blocks. Jason Helge Anderson, Stephen Brown 0003 |
FPGA | 2 |
| 1997 | On two-step routing for FPGASabstractWe present results which show that a separate global and detailed routing strategy can be competitive with a combined routing process.Under restricted architectural assumptions, we compute a new lower bound for detailed routing and show that our detailed router typically requires no more than two extra routing tracks above this computed limit.Also, experimental results show that the Mapping Anomaly presented in [20], which suggests that separated routing may yield arbitrarily poor results in certain instances, is a concern only if nets are restricted to a single track domain.Finally, to motivate future work, we show the latest two-step routing results that we have achieved with the VPR global router and SEGA detailed router tools on the largest CBL benchmark circuits.j Penllis&m to m&e digitnl/hnrd copies ofall or port of this materin for personal or chwoom use is granted witbout fee provided that the copies are not made or distributed for profit or commercial advantage.the COPY-._.-_._-. .~ Guy Lemieux, Stephen Brown 0003, Daniel Vranesic |
ISPD | 2 |
| 1996 | Experience in Designing a Large-scale Multiprocessor using Field-Programmable Devices and Advanced CAD ToolsabstractThis paper provides a case study that shows how a demanding application stresses the capabilities of today's CAD tools, especially in the integration of products from multiple vendors.We relate our experiences in the design of a large, high-speed multiprocessor computer, using state of the art CAD tools.All logic circuitry is targeted to field-programmable devices (FPDs).This choice amplifies the difficulties associated with achieving a highspeed design, and places extra requirements on the CAD tools.Two main CAD systems are discussed in the paper: Cadence Logic Workbench (LWB) is employed for board-level design, and Altera MAX+plusII is used for implementation of logic circuits in FPDs.Each of these products is of great value for our project, but the integration of the two is less than satisfactory.The paper describes a custom procedure that we developed for integrating sub-designs realized in FPDs (via MAX+plusII) into our board-level designs in LWB.We also discuss experiences with Logic Modelling Smart Models, for simulation of FPDs and other types of chips. Stephen Brown 0003, Naraig Manjikian, Zvonko G. Vranesic, S. Caranci, A. Grbic, R. Grindley, M. Gusat, K. Loveless, Zeljko Zilic, Sinisa Srbljic |
DAC | 1 |
| 1996 | Hybrid FPGA ArchitectureabstractNo abstract available. Alireza Kaviani, Stephen Brown 0003 |
FPGA | 2 |
| 1993 | A stochastic model to predict the routability of field-programmable gate arraysabstractOne area of particular importance is the design of an FPGA routing architecture, which houses the user-programmable switches and wires that are used to interconnect the FPGAs logic resources. Because the routing switches consume significant chip area and introduce propagation delays, the design of the routing architecture greatly influences both the area utilization and speed performance of an FPGA. FPGA routing architectures have already been studied using experimental techniques. This paper describes a stochastic model that facilitates exploration of a wide range of FPGA routing architectures using a theoretical approach. In the stochastic model an FPGA is represented as an N*N array of logic blocks separated by both horizontal and vertical routing channels, similar to a Xilinx FPGA. A circuit to be routed is represented by additional parameters that specify the total number of connections, and each connection's length and trajectory. The stochastic model gives an analytic expression for the routability of the circuit in the FPGA. Practically speaking, routability can be viewed as the likelihood that a circuit can be successfully routed in a given FPGA. The routability predictions from the model are validated by comparing them with the results of a previously published experimental study on FPGA routability.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1992 | Improving FPGA Routing Architectures Using Architecture and CAD InteractionsabstractThe interactions between the CAD tools that are used to configure the routing resources of a field-programmable gate array (FPGA) and the design of the routing architecture itself are examined. Such an understanding is used to determine where to reduce the number of routing switches in the FPGA while maintaining routability. Experiments are used to study a switch block that was previously thought to have unacceptably low flexibility. It is shown that the performance of this switch block can be improved by adapting the global router to require less flexibility in the architecture, and by careful placement of physical pins on the logic blocks. It is demonstrated that the fewest routing switches are required when each logical pin appears on only one side of the logic cell rather than two or more.> Benjamin Tseng, Jonathan Rose, Stephen Brown 0003 |
ICCD | 3 |
| 1992 | A detailed router for field-programmable gate arraysabstractA detailed routing algorithm, called the coarse graph expander (CGE), that has been designed specifically for field-programmable gate arrays (FPGAs) is described. The algorithm approaches this problem in a general way, allowing it to be used over a wide range of different FPGA routing architectures. It addresses the issue of scarce routing resources by considering the side effects that the routing of one connection has on another, and also has the ability to optimize the routing delays of time-critical connections. CGE has been used to obtain excellent routing results for several industrial circuits implemented in FPGAs with various routing architectures. The results show that CGE can route relatively large FPGAs in very close to the minimum number of tracks as determined by global routing, and it can successfully optimize the routing delays of time-critical connections. CGE has a linear run time over circuit size.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1990 | A Detailed Router for Field-Programmable Gate ArraysabstractThe course graph expansion (CGE) detailed routing algorithm is presented for FPGAs (field-programmable gate arrays). The algorithm has the ability to resolve routing conflicts by considering the side-effects of one connection on another, and can be used over a wide range of FPGA interconnection architectures. CGE has been used to obtain excellent routing results for several industrial circuits with various FPGA routing architectures. The results show that CGE is able to route relatively large FPGAs in the absolute minimum number of tracks as determined by global routing, and that CGE has a linear run-time over circuit size.> Stephen Brown 0003, Jonathan Rose, Zvonko G. Vranesic |
ICCAD | 1 |