EDBT 2026 Demo / reviewers in the wild / expert
Herman Schmit
dblp:s/HermanSchmit
· DBLP profile ↗
41ranked-venue papers
16as first author
3since 2021 · last 2023
0000-0002-0109-7604ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 16 first-author · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Improving Standard-Cell Design Flow using Factored Form OptimizationabstractFactored form is a powerful multi-level representation of a Boolean function that readily translates into an implementation of the function in CMOS technology. In particular, the number of literals in a factored form correlates strongly with the number of transistors in the CMOS implementation. This paper develops novel methods for optimizing factored forms while working on the efficient and-inverter graph (AIG) representation of combinational logic. This is in contrast to the traditional logic synthesis based on logic networks, and other AIG-based methods that minimize the AIG nodes count. Experiments show that applying these methods helps to reduce the area after technology mapping by an additional 2.8% on average, compared to a high-effort area-oriented baseline. It is expected that deploying these methods as part of an industrial standard-cell design flow will reduce design costs and power consumption. Additionally, this work enables efficient transistor-level logic synthesis of large designs with various applications in design automation. Alessandro Tempia Calvino, Alan Mishchenko, Herman Schmit, Ethan Mahintorabi, Giovanni De Micheli |
DAC | 3 |
| 2022 | Multi-input Serial Adders for FPGA-like Computational FabricabstractIn this paper, we present a new functional unit to replace the LUT in an FPGA-like computational fabric designed specifically for use to accelerate instance-specific sparse integer matrix multiplication. We use a suite of matrices, the VPR place-and-route tool, and modern architecture representations of the interconnect to examine this architectural idea. The new cell, called the K--ADD, increases density by 2.5x to 4x, and increases performance by 8% to 30% by simultaneously increasing the clock rate and reducing the number of cycles to compute the product. This benefit magnifies the two-orders-of-magnitude advantage of using instance-specific matrix multipliers demonstrated in prior work. We investigate the cluster size, N, across multiple technology nodes. In that investigation, we see a sustained benefit to a larger cluster size (N=8). This observation holds for both netlists mapped to a 6--LUT and to a 6--ADD, which implies this behavior has more to do with the peculiar structure of these matrix multiplication netlists, not the different functional unit. Herman Schmit, Matthew Denton |
FPGA | 1 |
| 2022 | Direct Spatial Implementation of Sparse Matrix Multipliers for Reservoir ComputingabstractReservoir computing is a nascent sub-field of machine learning which relies on the recurrent multiplication of a very large, sparse, fixed matrix. We argue that direct spatial implementation of these fixed matrices minimizes the work performed in the computation, and allows for significant reduction in latency and power through constant propagation and logic minimization. Bit-serial arithmetic enables massive static matrices to be implemented. We present the structure of our bit-serial matrix multiplier, and evaluate using canonical signed digit representation to further reduce logic utilization. We have implemented these matrices on a large FPGA and provide a cost model that is simple and extensible. These FPGA implementations, on average, reduce latency by 50x up to 86x versus GPU libraries. Comparing against a recent sparse DNN accelerator, we measure a 4.1x to 47x reduction in latency depending on matrix dimension and sparsity. Matthew Denton, Herman Schmit |
HPCA | 2 |
| 2019 | Spatial Timing Analysis With Exact Propagation of Delay and Application to FPGA PerformanceabstractThis paper introduces a method for delay approximation in timing analysis for spatial variation where the parameters representing the sources of variation are based on a convolution of random variation with a basis function. This results in delay models that need only include nine nonzero random variables for each edge in a typical spatial correlation model, less than previous work. As a result the symbolic representation of all potentially critical delays can be carried through the entire timing analysis without the need for approximate maximum function of statistical delays, producing a conservative but accurate upper bound on timing to meet a specified yield target. The algorithm has been applied to a set of field-programmable gate array designs and used to show that it is computationally efficient even with large designs. Further, because the timing is pessimistic but accurate, it removes the need for guardbanding the timing analysis compared to previous approaches. The representation of all potentially critical paths in a symbolic form also enables fast Monte-Carlo (MC) algorithms to supplement the deterministic ones. We show a combined deterministic and MC algorithm that is on average within 0.35% of ideal, and has a 10-6probability of an optimistic result. David M. Lewis, Herman Schmit |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Stratix™ 10 High Performance Routable Clock NetworksabstractWe present the clock architecture of the Stratix?10 FPGA, which uses a routable clock network rather than the fixed clock networks of previous generations. We describe the flexibility provided by this routable clock network and how arbitrarily sized clock trees can be synthesized and placed anywhere on the FPGA. We show how this capability to generate customized clock trees can provide better performance through reduced clock loss while maintaining the ability to handle the large number of clock domains that modern systems require. We experimentally demonstrate how a routable clock tree reduces the clock loss of the user design implementation by up to 6% of clock insertion delay. Carl Ebeling, Dana How, David M. Lewis, Herman Schmit |
FPGA | 4 |
| 2016 | Dissecting Xeon + FPGA: Why the integration of CPUs and FPGAs makes a power difference for the datacenter: Invited PaperabstractIntel's Xeon roadmap includes package-integrated FPGAs in every new generation. In this talk, we will dissect why this is such a powerful combination at this time of great change in datacenter workloads. We will show how power savings within the CPU complex is a significant multiplier for power savings in the datacenter as a whole. Focusing on the domain of machine learning, we will present the recent evolution of data types and operators, and make the case that FPGAs are the path to facilitate this continued evolution. Finally, we will discuss the criticality of the close coupling of the CPU and the FPGA. This coupling facilitates high bandwidth and low latency communication that is required for the development, debugging and deployment of heterogeneous applications. Herman Schmit, Randy Huang |
ISLPED | 1 |
| 2008 | Placement challenges for structured ASICsabstractThe placement problem for structured ASICs combines aspects of the standard cell ASIC placement problem and FPGA placement. Similarities with ASIC placement include the number and size of the place-able objects and the need to consider buffering within placement. Similarities with FPGA placement include the existence of discrete legal locations for all types of objects, the constraints caused by "intrinsic" connections, such as clock, reset or IO signals and fixed routing tracks. The research community has provided detailed analysis of various different solutions for the standard cell placement problem over the last two decades. FPGA placement research has not focused on the legalization issues. Architecturally, FPGAs are changing to focus more on synthesis and clustering than fine-grained placement to meet timing. In this paper we discuss the similarities and differences between FPGA, Standard Cell, and Structured ASIC placement, and we present new representations and tests cases for the structured ASIC problem Herman Schmit, Radu Ciobanu |
ISPD | 1 |
| 2005 | Layout techniques for FPGA switch blocksabstractThis paper presents abstract layout techniques for a variety of field-programmable gate array switch block architectures. For subset switch blocks of small size, we find the optimal implementations using a simple metric. We also develop a tractable heuristic that returns the optimal results for small switch blocks and good results for large switch blocks. We show how it is possible to transform universal switch blocks into a subset architecture by using the decomposition property of universal switch blocks. This allows universal switch blocks to exploit the same layout methodologies as presented for subset architectures. Herman Schmit, Vikas Chandra |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Enabling energy efficiency in via-patterned gate array devicesabstractIn an attempt to enable the cost-effective production of low-and mid-volume application-specific chips, researchers have proposed a number of so-called structured ASIC architectures. These architectures represent a departure from traditional standard-cell-based ASIC designs in favor of techniques which present more physical and structural regularity. If structured ASICs are to become a viable alternative to standard cells, they must deliver performance and energy efficiency which is competitive with standard-cell-based design techniques. This paper focuses on one family of structured ASICs known as via-patterned gate arrays, or VPGAs. In this paper, we present circuit structures and power optimization algorithms which can be applied to VPGA chips in an effort to reduce their operational power dissipation. R. Reed Taylor, Herman Schmit |
DAC | 2 |
| 2004 | An Interconnect Channel Design Methodology for High Performance Integrated CircuitsabstractOn-chip communication is becoming a bottleneck for high performance designs. Conventional interconnect design methodology does not account for architectures and/or communication schemes that require storage buffers (first-in-first-out queues or FIFOs) in the interconnect channel. For example, FIFOs and flow-control are needed for Network-on-Chip, high performance ASICs and multiple clock domain designs. These IC implementation architectures require an efficient methodology to determine the size of the FIFOs in the channel since the FIFO sizes affect system performance. In this work we devised a methodology to size the FIFOs in an interconnect channel containing one or more FIFOs connected in series. We show that the sizing of the FIFOs in the channel is a function of system parameters such as data production rate and consumption rate, data burstiness, number of channel stages etc. and we also quantify their effect on performance. For a single clock design, we have developed an efficient algorithm which reduces the search space for the optimal sizing of the FIFOs in the channel. Vikas Chandra, Anthony Xu, Herman Schmit, Lawrence T. Pileggi |
DATE | 3 |
| 2004 | A power aware system level interconnect design methodology for latency-insensitive systemsabstractLatency-insensitive interconnects require first-in-first-out buffers (FIFO) for flow-control and storage. Interconnect delays are not scaling in proportion to the clock period and hence multiple stages of FIFOs will be needed for high performance interconnects. FIFOs in the interconnect are a significant contributor to the total power consumption. In this work, we propose a design methodology to synthesize a low power interconnect channel containing series connected FIFOs for latency-insensitive systems. Our approach is the first to consider and simultaneously optimize the channel clock frequency, voltage and the FIFO sizes to minimize the power consumption. For small problem size, we show that our approach finds solutions which are close to optimal. The power aware interconnect channel synthesis is affected by the system parameters like the data production rate and data consumption rate. The choice of optimal channel clock frequency, voltage and FIFO sizes can lead to power savings as high as 77.7%, 83.6% and 87% for a 3 stage, 4 stage and a 5 stage channel respectively. Vikas Chandra, Herman Schmit, Anthony Xu, Lawrence T. Pileggi |
ICCAD | 2 |
| 2004 | Creating a power-aware structured ASICabstractIn an attempt to enable the cost-effective production of low-and mid-volume application-specific chips, researchers have proposed a number of so-called structured ASIC architectures. These architectures represent a departure from traditional standard-cell-based ASIC designs in favor of techniques which present more physical and structural regularity. This paper presents circuits which provide power-performance exibility in this regular, structured ASIC environment. These circuits, which employ gate sizing and voltage scaling for energy efficiency, enable delay-constrained power optimization to be performed for structured ASIC designs. R. Reed Taylor, Herman Schmit |
ISLPED | 2 |
| 2003 | Exploring regular fabrics to optimize the performance-cost trade-offabstractWhile advances in semiconductor technologies have pushed achievable scale and performance to phenomenal limits for ICs, nanoscale physical realities dictate IC production based on what we can afford. We believe that IC design and manufacturing can be made more affordable, and reliable, by removing some design and implementation flexibility and enforcing new forms of design regularity. This paper discusses some of the trade-offs to consider for determination of how much regularity a particular IC or application can afford. A Via Patterned Gate Array is proposed as one such example that trades performance for cost by way of new forms of design regularity. Lawrence T. Pileggi, Herman Schmit, Andrzej J. Strojwas, Padmini Gopalakrishnan, V. Kheterpal, Aneesh Koorapaty, Chetan Patel, Vyacheslav Rovner, Kim Yaw Tong |
DAC | 2 |
| 2003 | Heterogeneous Programmable Logic Block Architectures
Aneesh Koorapaty, Vikas Chandra, Kim Yaw Tong, Chetan Patel, Lawrence T. Pileggi, Herman Schmit |
DATE | 6 |
| 2003 | Asynchronous PipeRench: Architecture and Performance EstimationsabstractPipeRench is a configurable architecture that has the unique ability to virtualize an application using dynamic reconfiguration. This paper investigates the potential benefits and costs of implementing this architecture using an asynchronous methodology. Since clock distribution and gating are relatively easy in the synchronous PipeRench, we focus on the benefit due to decreased timing pessimism in an asynchronous implementation. Two architectures for fully asynchronous implementation are considered. PE-based asynchronous implementation yields approximately 80% improvement in performance per stripe. This implementation, however, requires significant increases in configuration storage and wire count. A few particular features of the architecture, such as the crossbar interconnect structure within the stripe, are primarily responsible for this growth in configuration bits and wires. These features, however, are the primary aspects of the PipeRench architecture that make it a good compilation target. Hiroto Kagotani, Herman Schmit |
FCCM | 2 |
| 2003 | Efficient Application Representation for HASTE: Hybrid Architectures with a Single, Transformable ExecutableabstractHybrid architectures, which are composed of a conventional processor closely coupled with reconfigurable logic, seem to combine the advantages of both types of hardware. They present some practical difficulties however. The interface between the processor and the reconfigurable logic is crucial to performance and is often difficult to implement well. Partitioning the application between the processor and logic is a difficult task, typically complicated by entirely different programming models, heterogeneous interfaces to external resources, and incompatible representations of applications. A separate executable must be produced and maintained for each type of hardware. An architecture called HASTE (Hybrid Architecture with a Single Transformable Executable) solves many of these difficulties. HASTE allows a single executable to represent an entire application, including portions that run on a reconfigurable fabric and portions that run on a sequential processor. This executable can execute in its entirety on the processor, but for best performance portions of the application that are mapped onto the fabric at run-time. The application representation is the key to making this concept viable, and several different ones were examined. Some used a relatively conventional register instruction set architecture (ISA) while others used a new queue-based ISA. AN ISA using a modified form of register addressing has been shown to have the best overall characteristics and should allow for the practical implementation of HASTE. Benjamin A. Levine, Herman Schmit |
FCCM | 2 |
| 2003 | Heterogeneous Logic Block Architectures for Via-Patterned Programmable Fabrics
Aneesh Koorapaty, Lawrence T. Pileggi, Herman Schmit |
FPL | 3 |
| 2003 | Extra-dimensional Island-Style FPGAs
Herman Schmit |
FPL | 1 |
| 2003 | Floorplanning of pipelined array modules using sequence pairsabstractFloorplanning individual pipelined array modules of a larger overall die can yield beneficial results. Critical paths in every pipeline stage of a pipelined design are roughly equivalent after synthesis. The inability of synthesis tools to predict without full placement both wire congestion and the distance traveled by a wire or wires between consecutive registers are the greatest causes of additional delay and area during place and route. This paper will detail a floorplanning methodology for pipelined arrays that is used to regulate wire congestion and the shortest/longest distances travelled by wire(s) between consecutive registers. A new wire length metric for pipelined arrays will be discussed that attempts to measure the distance travelled by wire(s) between registers. A new move set for floorplanning pipelined arrays using sequence pairs will also be introduced that significantly reduces the annealing design space from previous work. These two contributions when used together have produced up to 10% faster clock periods, 12% smaller designs, and 85% less area used to fix hold time violations in a placed and routed 0.18 μm design. Matthew Moe, Herman Schmit |
ISPD | 2 |
| 2003 | An architectural exploration of via patterned gate arraysabstractIn this work we investigate the architecture of a Via Patterned Gate Array (VPGA) [1], focusing primarily on: 1) the optimal lookup table (LUT) size; and 2) a comparison the crossbar and switch block routing architectures. Unlike FPGAs, the routing architectures in a VPGA do not dominate the total area of the circuit. Therefore our results suggest that using smaller LUTs results in a much faster and smaller design. In the routing architecture comparison, our results also show that the switch block architecture is inferior to the crossbar architecture in terms of area utilization. As the number of routing tracks grows, the switch block architecture begins to dominate the total area of the design as in the case of the FPGAs. Chetan Patel, Anthony Cozzie, Herman Schmit, Lawrence T. Pileggi |
ISPD | 3 |
| 2002 | Memory optimization in single chip network switch fabricsabstractMoving high bandwidth (10Gb/s+) network switches from the large scale, rack mount design space to the single chip design space requires a re-evaluation of the overall design requirements. In this paper, we explore the design space for these single chip devices by evaluating the ITRS. We find that unlike ten years ago when interconnect was scarce, the limiting factor in today's designs is on-chip memory. We then discuss an architectural technique for maximizing the effectiveness of queue memory in a single chip switch. Next, we show simulation results that indicate that a more than two order of magnitude improvement in dropped packet probability can be achieved by re-distributing memory and allowing sharing between the switch's ports. Finally, we evaluate the cost of the optimized architecture in terms of other on-chip resources. David Whelihan, Herman Schmit |
DAC | 2 |
| 2002 | Queue Machines: Hardware Compilation in HardwareabstractIn this paper we hypothesize that reconfigurable computing is not more widely used because of the logistical difficulties caused by the close coupling of applications and hardware platforms. As an alternative, we propose computing machines that use a single, serial instruction representation for the entire reconfigurable computing application. We show how it is possible to convert, at runtime, the parallel portions of the application into a spatial representation suitable for execution on a reconfigurable fabric. The conversion to spatial representation is facilitated by the use of an instruction set architecture based on an operand queue. We describe techniques to generate code for queue machines and hardware virtualization techniques necessary to allow any application to execute on any platform. Herman Schmit, Benjamin A. Levine, Benjamin Ylvisaker |
FCCM | 1 |
| 2002 | FPGA switch block layout and evaluationabstractThis paper presents abstract layout techniques for a variety of FPGA switch block architectures. We evaluate the relative density of subset, universal, and Wilton switch block architectures. For subset switch blocks of small size, we find the optimal implementations using a simple metric. We also develop a tractable heuristic that returns the optimal results for small switch blocks, and good results for large switch blocks. For switch blocks with general connectivity, we develop a representation and a layout evaluation technique. We use these techniques to compare a variety of small switch blocks. We find that the traditional Xilinx-style, subset switch block is superior to the other proposed architectures. Finally, we have hand-designed some small switch blocks to confirm our results. Herman Schmit, Vikas Chandra |
FPGA | 1 |
| 2002 | Morphable Multipliers
Silviu M. S. A. Chiricescu, Michael A. Schuette, Robin Glinton, Herman Schmit |
FPL | 4 |
| 2000 | Implementation of Near Shannon Limit Error-Correcting Codes Using Reconfigurable HardwareabstractError correcting codes (ECCs) are widely used in digital communications. New types of ECCs have been proposed which permit error-free data transmission over noisy channels at rates which approach the Shannon capacity. For wireless communication, these new codes allow more data to be carried in the same spectrum, lower transmission power, and higher data security and compression. One new type of ECC, referred to as Turbo Codes, has received a lot of attention, but is computationally expensive to decode and difficult to realize in hardware. Low density parity check codes (LDPCs), another ECC, also provide near Shannon limit error correction ability. However, LDPCs use a decoding scheme which is much more amenable to hardware implementation. This paper first presents an overview of these coding schemes, then discusses the issues involved in building an LDPC decoder using reconfigurable hardware. It presents a hypothetical LDPC implementation using a commercial FPGA, which will give an idea of future research issues and performance gains. Benjamin A. Levine, R. Reed Taylor, Herman Schmit |
FCCM | 3 |
| 2000 | The John Henry Syndrome (panel session)(abstract only): humans vs. machines as FPGA designersabstractHuman designers have done amazing things with FPGAs. These designs challenge our assumptions about the speeds and densities acheivable by programmable hardware. But with multi-million gate designs and increasingly complex FPGA architectures is there really any place for the hand-crafted design anymore? Is there a way that CAD tools can incorporate the techniques and knowledge of designers to create high-density, high-performance implementations automatically? Or will the tools and architectures always lag the applications, thereby guaranteeing abundant job opportunities for FPGA design experts? Herman Schmit, Ray Andraka, Philip Friedin, Satnam Singh, Tim Southgate |
FPGA | 1 |
| 2000 | Scalable interconnect and power distribution for island-style FPGAs (poster abstract)abstractNo abstract available. Herman Schmit, David Whelihan, Peter Kamarchik, Frank Gennari |
FPGA | 1 |
| 2000 | PipeRench implementation of the instruction path coprocessorabstractThe paper demonstrates how an Instruction Path Coprocessor (I-COP) can be efficiently implemented using the PipeRench reconfigurable architecture. An I-COP is a programmable on-chip coprocessor that operates on the core processor's instructions to transform them into a new format that can be more efficiently executed. The I-COP can be used to implement many sophisticated hardware code modification techniques. We show how four specific techniques can be mapped to the PipeRench pipelined computation model. The experimental results show that a PipeRench I-COP used to perform trace construction and trace optimizations for a trace cache fill unit not only achieves good performance gains but can potentially be implemented in less than 10 mm/sup 2/ (assuming 0.18 micron technology) or approximately 3% of the die area of a current high-end microprocessor. We believe these results demonstrate the usefulness and feasibility of the I-COP concept. Yuan C. Chou, Pazhani Pillai, Herman Schmit, John Paul Shen |
MICRO | 3 |
| 1999 | Vertical Benchmarks for CADabstractVertical benchmarks are complex system designs represented at multiple levels of abstraction. More effective than componentbased CAD benchmarks, vertical benchmarks enable quantitative comparison of CAD techniques within or across design flows. This work describes the notion of vertical benchmarks and presents our benchmark, which is based on a commercial DSP, by comparing two alternative design flows. 2. Christopher Inacio, Herman Schmit, David Nagle, Andrew Ryan, Donald E. Thomas, Yingfai Tong, Ben Klass |
DAC | 2 |
| 1999 | PCI-PipeRench and the SWORDAPI: A System for Stream-Based Reconfigurable ComputingabstractReconfigurable hardware accelerators have been shown to be flexible and efficient in stream-based applications. In this paper, we discuss the design of PCI-PipeRench and the SWORDAPI. PCI-PipeRench is a coprocessor utilizing the PipeRench architecture which includes on-chip control and data buffering to interface with a host processor over a PCI bus. SWORDAPI calls resemble standard C file control functions, and allow developers to create applications Independent of underlying reconfigurable hardware details. In addition, the SWORDAPI provides a cosimulation environment so that verification can be performed using unmodified application source with a hardware simulator. Efficient utilization of the bus is of critical importance in the design of such a system; various methods used to address this issue are presented. Ronald Laufer, R. Reed Taylor, Herman Schmit |
FCCM | 3 |
| 1999 | Extra-Dimensional Island-Style FPGAsabstractNo abstract available. Herman Schmit |
FPGA | 1 |
| 1999 | PipeRench: A Coprocessor for Streaming multimedia AccelerationabstractFuture computing workloads will emphasize an architecture's ability to perform relatively simple calculations on massive quantities of mixed-width data. This paper describes a novel reconfigurable fabric architecture, PipeRench, optimized to accelerate these types of computations. PipeRench enables fast, robust compilers, supports forward compatibility, and virtualizes configurations, thus removing the fixed size constraint present in other fabrics. For the first time we explore how the bit-width of processing elements affects performance and show how the PipeRench architecture has been optimized to balance the needs of the compiler against the realities of silicon. Finally, we demonstrate extreme performance speedup on certain computing kernels (up to 190x versus a modern RISC processor), and analyze how this acceleration translates to application speedup. Seth Copen Goldstein, Herman Schmit, Matthew Moe, Mihai Budiu, Srihari Cadambi, R. Reed Taylor, Ronald Laufer |
ISCA | 2 |
| 1999 | Mixed-swing quadrail for low power dual-rail domino logicabstractThis paper describes a new mixed-swing topology for dual-rail domino logic that results in a simultaneous energy and delay reduction. HSPICE simulation results for a 1-bit full adder cell show a 24 % delay decrease and a 24 % energy reduction for the mixed-swing topology compared to standard dual-rail domino. Energy and delay trends with supply voltage scaling are also presented for the adder cell. An 8-bit by 8-bit multiplier design with mixedswing dual-rail domino adders is presented. Simulation results show this implementation to be 10 % faster with an 18 % energy savings. 1. Bharath Ramasubramanian, Herman Schmit, L. Richard Carley |
ISLPED | 2 |
| 1998 | Characterization and Parameterization of a Pipeline Reconfigurable FPGAabstractThe article defines a class of architectures for pipeline reconfigurable FPGAs by parameterizing a generic model. This class of architecture is sufficiently general to allow exploration of the most important design trade-offs. The parameters include the word size and LUT size, the number of global busses and registers associated with each logic block, and the horizontal interconnect within each stripe. We have developed an area model for the architecture that allows us to quickly estimate the area of an instance of the architectural class as a function of the parameter values. We compare the estimates generated by this model to one instance of the architecture that we have designed and fabricated. Matthew Moe, Herman Schmit, Seth Copen Goldstein |
FCCM | 2 |
| 1998 | Managing Pipeline-Reconfigurable FPGAsabstractWhile reconfigurable computing promises to deliver incomparable performance, it is still a marginal technology due to the high cost of developing and upgrading applications. Hardware virtualization can be used to significantly reduce both these costs. In this paper we describe the benefits of hardware virtualization, and show how it can be achieved using a combination of pipeline reconfiguration and run-time scheduling of both configuration streams and data streams. The result is PipeRench, an architecture that supports robust compilation and provides forward compatibility. Our preliminary performance analysis predicts that PipeRench will outperform commercial FPGAs and DSPs in both overall performance and in performance per mm2. Srihari Cadambi, Jeffrey Weener, Seth Copen Goldstein, Herman Schmit, Donald E. Thomas |
FPGA | 4 |
| 1998 | Address generation for memories containing multiple arraysabstractWe present techniques for generating addresses for memories containing multiple arrays. Because these techniques rely on the inversion or rearrangement of address bits, they are faster and require less hardware to compute than the traditional technique of addition. Use of these techniques can improve performance and cost of application-specific memory subsystems by decreasing effective access time to arrays and by reducing address generation hardware. The primary drawback to this approach is that extra memory space is occasionally required, but in over a million tested cases, this extra memory space is on average only 2% and no worse than 17.4% of the utilized memory space. This amount of wasted address space is significantly less than the amount required by the only known similar technique and rarely necessitates the allocation of additional memory components. These techniques provide a foundation for adder-free address generation for manually and automatically generated application-specific memory designs. Herman Schmit, Donald E. Thomas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1997 | Incremental reconfiguration for pipelined applicationsabstractThis paper examines the implementation of pipelined applications using run-time reconfiguration. Throughput and latency of pipelined applications can be significantly improved when reconfiguration is performed at the level of individual pipeline stages, as opposed to configuration of the entire FPGA. If reconfiguration and execution can be performed simultaneously, the performance of a pipelined application approaches its theoretical maximum. This paper proposes a new FPGA configuration mechanism, called striping, that supports pipeline stage reconfiguration and simultaneous configuration and execution. Additionally, the use of the pipeline stage as the atomic unit of reconfiguration introduces a design abstraction that enables the development families of upwardly-compatible FPGAs and virtual hardware design. Herman Schmit |
FCCM | 1 |
| 1997 | Is Reconfigurable Computing Commercially Viable (panel)?abstractArticle Is reconfigurable computing commercially viable (panel)? Share on Author: Herman Schmit Carnegie Mellon Univ. Carnegie Mellon Univ.View Profile Authors Info & Claims FPGA '97: Proceedings of the 1997 ACM fifth international symposium on Field-programmable gate arraysFebruary 1997 https://doi.org/10.1145/258305.258318Online:09 February 1997Publication History 0citation263DownloadsMetricsTotal Citations0Total Downloads263Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Herman Schmit |
FPGA | 1 |
| 1997 | Synthesis of application-specific memory designsabstractThis paper discusses the mapping of arrays in a behavior to memories in an implementation. We introduce a novel approach to the design of memory systems, which is based on a variety of array grouping techniques and dimensional transformations, and the binding of array groups to memory components with different dimensions, access times, and number of ports. The results of design actions are computed in terms of memory cost, the number of wires necessary to connect the memory to the data path, and the limit of performance imposed by the memory design on the implementation. Three different procedures can be used to find a suitable memory design. All three procedures are directed by a weighted and constrained system cost function, which enables the expression of the user's design priorities. Compared to related research efforts, our approach improves performance by as much as 19%, reduces memory cost as 40%, and decreases the number of wires required to connect the memory to the data path by up to 57%. Herman Schmit, Donald E. Thomas |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1995 | Hidden Markov modeling and fuzzy controllers in FPGAsabstractThis paper compares software and FPGA-based hardware implementations of two applications. The first application uses hidden Markov models, and the second application is a fuzzy controller. Hidden Markov modeling is used for temporal pattern recognition and speech recognition in particular. Both applications are accelerated when implemented in FPGA-based hardware, but this acceleration is obtained by using different algorithms than those used in software implementations. These different algorithms produce slightly different outputs; therefore both solution quality and performance must be evaluated to compare hardware and software implementations. The experience of designing these applications has implications for hardware/software codesign tools and for the migration of existing software applications to FPGA-based hardware. Herman Schmit, Donald E. Thomas |
FCCM | 1 |
| 1995 | Address generation for memories containing multiple arraysabstractThis paper presents techniques for generating addresses for memories containing multiple arrays. Because these techniques rely on the inversion or rearrangement of address bits, they are faster and require less hardware to compute than offset addition. Use of these techniques can decrease effective access time to arrays and reduce address generation hardware. The primary drawback is that extra memory space is occasionally required by these techniques, but this extra memory space is on average only 4% and no worse than 25.2% of the utilized memory space. This amount of wasted address space is less than the amount required by similar techniques. Herman Schmit, Donald E. Thomas |
ICCAD | 1 |