Earl E. Swartzlander Jr.

dblp:35/4911 · DBLP profile ↗
← Back
118ranked-venue papers
25as first author
1since 2021 · last 2023
0000-0002-8699-5277ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 84 · 16 first-authorTheory of computation · 18 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 3Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2023 Improved Montgomery Multiplication
abstract
The Montgomery multiplication algorithm is used to perform multiplication coupled with modular reduction without the need to employ a division operation. Serial Montgomery implementations may operate at a bit, digit, or word level. An established classification scheme for serial implementations considers two dimensions: whether the multiplication and reduction computations are separated or integrated, and whether operand or product words are prioritized for scanning. Presented here is an augmented version of the taxonomy, which adds a third dimension. The new dimension characterizes the degree of parallelism in performing low-level (bit or digit) computations. Introducing a small degree of bit or digit level parallelism to a formerly serial approach can enhance performance, both through the typical benefit of parallel computation, and by opportunistically avoiding unnecessary computations. This can be achieved for a modest incremental cost in increased area in lieu of the larger cost of realizing a fully parallel solution. In this way, a new region on the latency-area curve becomes available for tradeoff considerations. The novel Rescheduled Montgomery Multiplier is presented as an experimental realization of the augmented taxonomy.
Trenton J. Grale, Earl E. Swartzlander Jr.
ARITH2
2019 Design and Analysis of Approximate Redundant Binary Multipliers
abstract
As technology scaling is reaching its limits, new approaches have been proposed for computional efficiency. Approximate computing is a promising technique for high performance and low power circuits as used in error-tolerant applications. Among approximate circuits, approximate arithmetic designs have attracted significant research interest. In this paper, the design of approximate redundant binary (RB) multipliers is studied. Two approximate Booth encoders and two RB 4:2 compressors based on RB (full and half) adders are proposed for the RB multipliers. The approximate design of the RB-Normal Binary (NB) converter in the RB multiplier is also studied by considering the error characteristics of both the approximate Booth encoders and the RB compressors. Both approximate and exact regular partial product arrays are used in the approximate RB multipliers to meet different accuracy requirements. Error analysis and hardware simulation results are provided. The proposed approximate RB multipliers are compared with previous approximate Booth multipliers; the results show that the approximate RB multipliers are better than approximate NB Booth multipliers especially when the word size is large. Case studies of error-resilient applications are also presented to show the validity of the proposed designs.
Weiqiang Liu 0001, Tian Cao 0005, Peipei Yin, Yuying Zhu 0003, Chenghua Wang, Earl E. Swartzlander Jr., Fabrizio Lombardi
IEEE Trans. Computers6
2017 High Performance Parallel Decimal Multipliers Using Hybrid BCD Codes
abstract
A parallel decimal multiplier with improved performance is proposed in this paper by exploiting the properties of three different binary coded decimal (BCD) codes, namely the redundant BCD excess-3 code (XS-3), the overloaded decimal digit set (ODDS) code and the BCD-4221/5211 code. The signed-digit radix-10 recoding is used to recode the BCD multiplier to the digit set [-5, 5] from [0, 9]. The redundant BCD XS-3 code is adopted to generate the multiplicand multiples in a carry-free manner. The XS-3 coded partial products (PPs) are converted to ODDS PPs to fit binary partial product reduction (PPR). In this paper, a regular decimal PPR tree using ODDS and BCD-4221/5211 codes is proposed; it consists of a binary PPR tree block, a non-fixed size BCD-4221 counter block and a BCD-4221/5211 PPR tree block. The decimal carry-save algorithm based on BCD-4221/5211 is used in the PPR tree to obtain high performance multipliers. Moreover, an improved PPG circuit and an improved parallel prefix/carry-select decimal adder are proposed to further improve the performance of the proposed multipliers. Analysis and comparison using the 45 nm technology show that the proposed decimal multipliers are faster and require less hardware area than previous designs found in the technical literature.
Xiao-Ping Cui, Wenwen Dong, Weiqiang Liu 0001, Earl E. Swartzlander Jr., Fabrizio Lombardi
IEEE Trans. Computers4
2016 A Modified Partial Product Generator for Redundant Binary Multipliers
abstract
Due to its high modularity and carry-free addition, a redundant binary (RB) representation can be used when designing high performance multipliers. The conventional RB multiplier requires an additional RB partial product (RBPP) row, because an error-correcting word (ECW) is generated by both the radix-4 Modified Booth encoding (MBE) and the RB encoding. This incurs in an additional RBPP accumulation stage for the MBE multiplier. In this paper, a new RB modified partial product generator (RBMPPG) is proposed; it removes the extra ECW and hence, it saves one RBPP accumulation stage. Therefore, the proposed RBMPPG generates fewer partial product rows than a conventional RB MBE multiplier. Simulation results show that the proposed RBMPPG based designs significantly improve the area and power consumption when the word length of each operand in the multiplier is at least 32 bits; these reductions over previous NB multiplier designs incur in a modest delay increase (approximately 5 percent). The power-delay product can be reduced by up to 59 percent using the proposed RB multipliers when compared with existing RB multipliers.
Xiao-Ping Cui, Weiqiang Liu 0001, Xin Chen 0039, Earl E. Swartzlander Jr., Fabrizio Lombardi
IEEE Trans. Computers4
2015 Low-Cost Duplicate Multiplication
abstract
Rising levels of integration, decreasing component reliabilities, and the ubiquity of computer systems make error protection a rising concern. Meanwhile, the uncertainty of future fault and error modes motivates the design of strong error detection mechanisms that offer fault-agnostic error protection. Current concurrent hardware mechanisms, however, either offer strong error detection coverage at high cost or restrict their coverage to narrow synthetic error models. This paper investigates the potential for duplication using alternate number systems to lower the costs of duplicated multiplication without sacrificing error coverage. Two examples of such low-cost duplication schemes are described and evaluated, it is shown that specialized carry-save or residue number system checking can be used to increase the efficiency of duplicated multiplication.
Michael B. Sullivan 0001, Earl E. Swartzlander Jr.
ARITH2
2014 Design of Goldschmidt Dividers with Quantum-Dot Cellular Automata
abstract
A Goldschmidt divider implemented with semiconductor quantum-dot cellular automata (QCA) is described. Most Goldschmidt dividers use a state machine for control, but state machines are difficult to implement in QCAs due to the long delays between the state machines and the computational circuits to be controlled. To resolve this problem, a data tag method is used. The data tags travel with the data and local tag decoders generate control signals. Since each datum has a tag, pipelining can be employed to increase the throughput. Using the new architecture, fixed-point Goldschmidt dividers are implemented with QCA technology.
Inwook Kong, Seong-Wan Kim, Earl E. Swartzlander Jr.
IEEE Trans. Computers3
2013 Improved Architectures for a Floating-Point Fused Dot Product Unit
abstract
This paper presents improved architectures for a floating-point fused two-term dot product unit. The floating-point fused dot product unit is useful for a wide variety of digital signal processing (DSP) applications including complex multiplication and fast Fourier transform (FFT) and discrete cosine transform (DCT) butterfly operations. In order to improve the performance, a new alignment scheme, early normalization, a four-input leading zero anticipation (LZA), a dual-path algorithm, and pipelining are applied. The proposed designs are implemented for single precision and synthesized with a 45nm standard cell library. The proposed dual-path design reduces the latency by 25% compared to the traditional floating-point fused dot product unit. Based on a data flow analysis, the proposed design can be split into three pipeline stages. Since the latencies of the three stages are fairly well balanced, the throughput is increased by a factor of 2.8 compared to the non-pipelined dual-path design.
Jongwook Sohn, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
2013 Truncated Logarithmic Approximation
abstract
The speed and levels of integration of modern devices have risen to the point that arithmetic can be performed very fast and with high precision. Precise arithmetic comes at a hidden cost-by computing results past the precision they require, systems inefficiently utilize their resources. Numerous designs over the past fifty years have demonstrated scalable efficiency by utilizing approximate logarithms. Many such designs are based off of a linear approximation algorithm developed by Mitchell. This paper evaluates a truncated form of binary logarithm as a replacement for Mitchell's algorithm. The truncated approximate logarithm simultaneously improves the efficiency and precision of Mitchell's approximation while remaining simple to implement.
Michael B. Sullivan 0001, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
2013 Fused floating-point two-term sum-of-squares unit
abstract
This paper proposes a low-power and high-performance fused floating-point two-term sum-of-squares (fused SoSQ) unit and compares it with discrete and fused dot-product units containing normal significand multipliers and discrete sum-of-squares units containing significand squarers; with these units, sum-of-squares is computed. The fused sum-of-squares unit has less latency, area, and power consumption than the floating-point dot-product units and the discrete parallel sum-of-squares unit. The main reason is that a fused floating-point architecture is used, its significand squarer is faster and smaller than normal significand multipliers, and the sub-modules related to subtraction can be removed. Furthermore, compound addition can be applied to the fused sum-of-squares unit to enhance the performance in the single-path addition. Compared with the fused floating-point dot-product, the fused floating-point sum-of-squares unit with the compound addition has 54% less power consumption, 48% less area, and 44% less latency.
Jae Hong Min, Earl E. Swartzlander Jr.
ASAP2
2013 Power analysis attack of QCA circuits: A case study of the Serpent cipher
abstract
Quantum-dot cellular automata (QCA) technology is an attractive alternative to CMOS for future digital designs. A powerful attack based on power analysis has become a significant threat to the security of CMOS cryptographic circuits. As there is no current flow in QCA, the power consumption of a QCA circuit is extremely low compared to its CMOS counterpart. Therefore, in this paper an investigation is carried out to ascertain if QCA circuits could be immune to power analysis attacks based on a case study of the Serpent cipher. In comparison to a previous design, the proposed QCA implementation of a sub-module of the Serpent cipher is more efficient in terms of complexity, area and latency. By using an upper bound power model, the first power analysis attack of a QCA cryptographic circuit is presented. Simulation results show that even though the power consumption is low, it can still be correlated with the correct key guess, and all possible subkeys applied to the Serpent sub-module can be revealed in a best case scenario for attackers. The security of practical QCA devices is also discussed and could be greatly improved by applying a smoother clock.
Weiqiang Liu 0001, Saket Srivastava, Máire O'Neill, Earl E. Swartzlander Jr.
ISCAS5
2013 STARS: Electronic Calculators: Desktop to Pocket
abstract
Discusses the historial development of computers from electronic calculators to current pocket-sized devices.
Earl E. Swartzlander Jr.
Proc. IEEE1
2013 QCA Systolic Array Design
abstract
Quantum-dot Cellular Automata (QCA) technology is a promising potential alternative to CMOS technology. To explore the characteristics of QCA and suitable design methodologies, digital circuit design approaches have been investigated. Due to the inherent wire delay in QCA, pipelined architectures appear to be a particularly suitable design technique. Also, because of the pipeline nature of QCA technology, it is not suitable for a complicated control system design. Systolic arrays take advantage of pipelining, parallelism, and simple local control. Therefore, an investigation into these architectures in semiconductor QCA technology is provided in this paper. Two case studies, (a matrix multiplier and a Galois Field multiplier) are designed and analyzed based on both multilayer and coplanar crossings. The performance of these two types of interconnections are compared and it is found that even though coplanar crossings are currently more practical, they tend to occupy a larger design area and incur slightly more delay. A general semiconductor QCA systolic array design methodology is also proposed. It is found that by applying a systolic array structure in QCA design, significant benefits can be achieved particularly with large systolic arrays, even more so than when applied in CMOS-based technology.
Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr.
IEEE Trans. Computers4
2013 Structure-Aware Placement Techniques for Designs With Datapaths
abstract
As technology scales and frequencies increase, a new hybrid design style emerges, wherein designs contain a mixture of random logic and datapath standard-cell components. This paper demonstrates that conventional half-perimeter wirelength driven placers underperform in terms of regularity and Steiner wirelength (StWL) for such hybrid designs. In addition, the quality gap between manual and automatic placement is more pronounced as the designs become more datapath oriented. To effectively handle hybrid designs, this paper proposes a new unified placement flow that simultaneously places random logic and datapath cells. This flow is built on the top of a leading academic force-directed placer and significantly improves the quality of datapath placement while leveraging the speed and flexibility of existing random-logic placement algorithms. It consists of a suite of novel global and detailed placement techniques, collectively called structure-aware placement techniques (SAPT). These techniques effectively integrate alignment constraints into placement, thereby overcoming the deficiencies of existing random-logic placers when handling designs with embedded datapaths. Compared to other state-of-the-art placers, SAPT improves total StWL by more than 28% and total routing overflow by over six times on the ISPD 2011 datapath benchmark suite. In addition, it improves total StWL by 5.8% on industrial hybrid designs.
Samuel I. Ward, Myung-Chul Kim, Natarajan Viswanathan, Zhuo Li 0001, Charles J. Alpert, Earl E. Swartzlander Jr., David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2012 Long Residue Checking for Adders
abstract
As system sizes grow and devices become more sensitive to faults, adder protection may be necessary to achieve system error-rate bounds. This study investigates a novel fault detection scheme for fast adders, long residue checking (LRC), which has substantive advantages over all previous separable approaches. Long residues are found to provide a ~10% reduction in complexity and ~25% reduction in power relative to the next most efficient error detector, while remaining modular and easy to implement.
Michael B. Sullivan 0001, Earl E. Swartzlander Jr.
ASAP2
2012 Cost-efficient decimal adder design in Quantum-dot cellular automata
abstract
Applications that cannot tolerate the loss of accuracy that results from binary arithmetic demand hardware decimal arithmetic designs. Binary arithmetic in Quantum-dot cellular automata (QCA) technology has been extensively investigated in recent years. However, only limited attention has been paid to QCA decimal arithmetic. In this paper, two cost-efficient binary-coded decimal (BCD) adders are presented. One is based on the carry flow adder (CFA) using a conventional correction method. The other uses the carry look ahead (CLA) algorithm which is the first QCA CLA decimal adder proposed to date. Compared with previous designs, both decimal adders achieve better performance in terms of latency and overall cost. The proposed CFA-based BCD adder has the smallest area with the least number of cells. The proposed CLA-based BCD adder is the fastest with an increase in speed of over 60% when compared with the previous fastest decimal QCA adder. It also has the lowest overall cost with a reduction of over 90% when compared with the previous most cost-efficient design.
Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr.
ISCAS4
2012 Keep it straight: teaching placement how to better handle designs with datapaths
abstract
As technology scales and frequency increases, a new design style is emerging, referred to as hybrid designs, which contain a mixture of random logic and datapath standard cell components. This work begins by demonstrating that conventional Half-Perimeter Wire Length (HPWL)-driven placers under-perform in terms of regularity and Steiner Wire Length (StWL) for such hybrid designs, and the quality gap between manual placement and automatic placers is more pronounced as the designs become more datapath-oriented. Then, a new unified placement flow that simultaneously handles random logic and datapath standard cells is proposed that significantly improves the placement quality of the datapath while leveraging the speed of modern state-of-the-art placement algorithms. The placement flow is built on top of a leading academic force-directed placer. It consists of a series of novel global and detailed placement techniques, collectively called Structure Aware Placement Techniques (SAPT). The techniques effectively integrate alignment constraints into placement, overcoming the deficiencies of the HPWL objective. Experimental results comparing our placement flow with six state-of-the-art placers on the ISPD 2011 Datapath Benchmark Suite show at least a 32% improvement in total StWL with over a 6x improvement in total routing overflow. In addition, the flow demonstrates an 8.25% improvement in total StWL on industrial hybrid designs.
Samuel I. Ward, Myung-Chul Kim, Natarajan Viswanathan, Zhuo Li 0001, Charles J. Alpert, Earl E. Swartzlander Jr., David Z. Pan
ISPD6
2012 A new hierarchical packet classification algorithm
Hyesook Lim, Earl E. Swartzlander Jr.
Comput. Networks3
2012 FFT Implementation with Fused Floating-Point Operations
abstract
This paper describes two fused floating-point operations and applies them to the implementation of fast Fourier transform (FFT) processors. The fused operations are a two-term dot product and an add-subtract unit. The FFT processors use "butterfly” operations that consist of multiplications, additions, and subtractions of complex valued data. Both radix-2 and radix-4 butterflies are implemented efficiently with the two fused floating-point operations. When placed and routed using a high performance standard cell technology, the fused FFT butterflies are about 15 percent faster and 30 percent smaller than a conventional implementation. Also the numerical results of the fused implementations are slightly more accurate, since they use fewer rounding operations.
Earl E. Swartzlander Jr., Hani Saleh
IEEE Trans. Computers1
2011 Design rules for Quantum-dot Cellular Automata
abstract
As a promising alternative to CMOS technology, QCA circuit design has been extensively studied in recent years. However, although a concrete set of design rules exist for integrated circuit design, little attention has been paid to the design rules necessary for efficient QCA circuit design. This paper compiles a set of important QCA design rules which include layout design rules, timing rules and some special rules for QCA technology to ensure QCA circuits function correctly and reliably. These rules will promote the development of practical and efficient QCA systems. A GF(2m) multiplier design is proposed as a case study to illustrate these design rules.
Weiqiang Liu 0001, Máire O'Neill, Earl E. Swartzlander Jr.
ISCAS4
2011 Quantifying academic placer performance on custom designs
abstract
There have been significant prior efforts to quantify performance of academic placement algorithms, primarily by creating artificial test cases that attempt to mimic real designs, such as the PEKO benchmark containing known optimas [5]. The idea was to create benchmarks with a known optimal solution and then measure how far existing placers were from the known optimal. Since the benchmarks do not necessarily correspond to properties of real VLSI netlists, the conclusions were met with some skepticism. This work presents two custom constructed datapath designs that perform common logic functions with hand-designed layouts for each. The new generation of academic placers is then compared against them to see how the placers performed for these design styles. Experiments show that all academic placers have wirelengths significantly greater then the manual solution; solutions range from 1.75 to 4.88 times greater wirelengths. These testcases will be released publically to stimulate research into automatically solving structured datapath placement problems.
Samuel I. Ward, David A. Papa, Zhuo Li 0001, Cliff C. N. Sze, Charles J. Alpert, Earl E. Swartzlander Jr.
ISPD6
2011 A Goldschmidt Division Method With Faster Than Quadratic Convergence
abstract
A new method to implement faster than quadratic convergence for Goldschmidt division using simple logic circuits is presented. While the approximate quotient converges quadratically in conventional Goldschmidt division, the new method achieves nearly cubic convergence. Although division with cubic convergence has been regarded as impractical due to its complexity, the proposed method reduces the logic complexity and the delay by using an approximate squarer with a simple logic implementation and a redundant binary Booth recoder. It is especially effective in a system that already has a radix-8 multiplier. As a result, the effective area for the reciprocal table can be reduced by 25.4%. The proposed method has been verified by SystemC and Verilog models. The final results are confirmed by simulation with both random double precision numbers and an exhaustive suite of 17-bit test vectors.
Inwook Kong, Earl E. Swartzlander Jr.
IEEE Trans. Very Large Scale Integr. Syst.2
2010 A novel technique for tunable mismatch shaping in oversampled digital-to-analog converters
abstract
Over-sampled digital-to-analog converters typically employ a unit-element architecture to drive out the analog signal. Performance can suffer from errors due to mismatch between unit elements, leading to a sharp drop in the achievable signal-to-noise ratio (SNR). Mismatch noise shaping is an established technique for overcoming these limitations, but usually anchors the signal band to a fixed location. In order to extend these advantages to tunable applications, this paper presents a novel technique that allows the mismatch noise shaping transfer function to have an adjustable center frequency.
Waqas Akram, Earl E. Swartzlander Jr.
ICASSP2
2010 A Rounding Method to Reduce the Required Multiplier Precision for Goldschmidt Division
abstract
A new rounding method to reduce the required precision of the multiplier for Goldschmidt division is presented. It applies special truncation methods at the final iteration step. This requires a minor modification to the rounding constants of the multiplier. It allows twice the error tolerance of conventional methods and inclusive error bounds. The proposed method further reduces the required precision of the multiplier by considering the asymmetric error bounds of Goldschmidt dividers where the factors are computed using a one's complement operation. As a result, the proposed rounding method allows the multiplier of a three-iteration Goldschmidt divider to be implemented using only three extra bits. The proposed method has been verified using a SystemC hardware model of the divider supporting variable precision. The validity of the error analysis is also checked via simulation. The final rounding results are checked with both 10^{10} random double precision floating-point significands and an exhaustive suite of 17-bit test vectors.
Inwook Kong, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2010 Priority Tries for IP Address Lookup
abstract
High-speed IP address lookup is essential to achieve wire speed packet forwarding in Internet routers. The longest prefix matching for IP address lookup is more complex than exact matching because it involves dual dimensions: length and value. This paper presents a new formulation for IP address lookup problem using range representation of prefixes and proposes an efficient binary trie structure named a priority trie. In this range representation, prefixes are represented as ranges on a number line between 0 and 1 without expanding to the maximum length. The best match to a given input address is the smallest range that includes the input. The priority trie is based on the trie structure, with empty internal nodes in the trie replaced by the priority prefix which is the longest among those in the subtrie rooted by the empty nodes. The search ends when an input matches a priority prefix, which significantly improves the search performance. Performance evaluation using real routing data shows that the proposed priority trie is very good in performance metrics such as lookup speed, memory size, update performance, and scalability.
Hyesook Lim, Changhoon Yim, Earl E. Swartzlander Jr.
IEEE Trans. Computers3
2010 Adaptive CORDIC: Using Parallel Angle Recoding to Accelerate Rotations
abstract
The CORDIC algorithm is used in the evaluation of a wide variety of elementary functions. It is a simple and elegant method, but it suffers from long latency. The Angle Recoding method is able to reduce the number of iterations by more than 50 percent, but its implementation in hardware requires a large increase in cycle time, to accommodate its complex angle selection function. This restricts its use to those cases where the angle of rotation is fixed and known in advance, so that the angle selection can be performed offline. This paper presents a simpler implementation of the angle selection scheme that does not require an increase in cycle time, thus allowing the Angle Recoding method to be used dynamically for arbitrary angles. The method also has the advantage that all the angle constants are found in parallel, in a single step, by testing only the initial rotation angle, without having to perform successive CORDIC iterations. This dynamic Angle Recoding method can be formulated to use ¿sections,¿ to limit the number of range comparators needed, to a reasonable value. There is an increase in the number of adaptive CORDIC iterations needed, but this problem can be mitigated by using a buffer in conjunction with the method of sections.
Terence K. Rodrigues, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2010 A Reduced Complexity Wallace Multiplier Reduction
abstract
Wallace high-speed multipliers use full adders and half adders in their reduction phase. Half adders do not reduce the number of partial product bits. Therefore, minimizing the number of half adders used in a multiplier reduction will reduce the complexity. A modification to the Wallace reduction is presented that ensures that the delay is the same as for the conventional Wallace reduction. The modified reduction method greatly reduces the number of half adders; producing implementations with 80 percent fewer half adders than standard Wallace multipliers, with a very slight increase in the number of full adders.
Ron S. Waters, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2009 A Power-Scalable Switch-Based Multi-processor FFT
abstract
This paper examines the architecture, algorithm and implementation of a switch-based multi-processor realization of the fast Fourier transform (FFT). The architecture employs M processing elements (PEs), and provides a speedup of M compared with systems that use a single PE. An algorithm is provided to detect and resolve memory conflicts. A CMOS implementation of a four-PE processor is presented. The design is reconfigurable to compute various FFT sizes. The design power consumption is scalable based on the number of active PEs. The timing, area and power results are discussed.
Bassam Jamil Mohd, Earl E. Swartzlander Jr.
ASAP2
2009 Adder and Multiplier Design in Quantum-Dot Cellular Automata
abstract
Quantum-dot cellular automata (QCA) is an emerging nanotechnology, with the potential for faster speed, smaller size, and lower power consumption than transistor-based technology. Quantum-dot cellular automata has a simple cell as the basic element. The cell is used as a building block to construct gates and wires. Previously, adder designs based on conventional designs were examined for implementation with QCA technology. That work demonstrated that the design trade-offs are very different in QCA. This paper utilizes the unique QCA characteristics to design a carry flow adder that is fast and efficient. Simulations indicate very attractive performance (i.e., complexity, area, and delay). This paper also explores the design of serial parallel multipliers. A serial parallel multiplier is designed and simulated with several different operand sizes.
Heumpil Cho, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2008 32 bit single cycle nonlinear VLSI cell for the ICA algorithm
abstract
The Independent Component Analysis (ICA) technique is amenable to a coarse-grain parallel-processing chip architecture. However, the computation of nonlinear functions is critical in this algorithm. An efficient hardware approach is presented here for the computation of such functions , some of which are compound and concatenated. All of the needed functions are regularized into a single algorithm so a new result produced on each cycle even if the function changes from one cycle to the next. The worst case arithmetic error is predicted and bounded. This enables the designer to quickly select the architectural parameters without expensive simulations, while insuring the desired accuracy. A design is presented for the 32 bit fixed point case.
Vijay K. Jain, Earl E. Swartzlander Jr.
ICASSP2
2008 A floating-point fused dot-product unit
abstract
A floating-point fused dot-product unit is presented that performs single-precision floating-point multiplication and addition operations on two pairs of data in a time that is only 150% the time required for a conventional floating-point multiplication. When placed and routed in a 45 nm process, the fused dot-product unit occupied about 70% of the area needed to implement a parallel dot-product unit using conventional floating-point adders and multipliers. The speed of the fused dot-product is 27% faster than the speed of the conventional parallel approach. The numerical result of the fused unit is more accurate because one rounding operation is needed versus at least three for other approaches.
Hani Saleh, Earl E. Swartzlander Jr.
ICCD2
2008 Speculative Carry Generation With Prefix Adder
abstract
A framework that generates formal prefix equations for speculative carry generation is presented. It is applicable to both normal carry and Ling carry adders and generates four forms of speculative carry generate prefix schemes (two forms for each carry case). For normal carry, one corresponds to an existing design and the other is newly introduced. For the Ling carry, both are newly proposed.
Youngmoon Choi, Earl E. Swartzlander Jr.
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Bridge Floating-Point Fused Multiply-Add Design
abstract
A new floating-point fused multiply-add (FMA) design for the execution of (A times B) + C as a single instruction is presented. The bridge fused multiply-add unit is a design intended to add FMA functionality to existing floating-point coprocessor units by including specialized hardware that reuses floating-point adder and floating-point multiplier components. The bridge unit adds this functionality without requiring an overhaul of coprocessor control units and without degrading the performance or parallel execution of addition and multiplication single instructions. To evaluate the performance, area, and power costs of adding a bridge FMA unit to common floating-point execution blocks, several circuits including a double-precision floating-point adder, floating-point multiplier, classic FMA, and a bridge FMA unit have been designed and implemented with AMD 65-nm silicon-on-insulator technology to provide a realistic and fair analysis of the presented FMA hardware tradeoffs.
Eric Quinnell, Earl E. Swartzlander Jr., Carl Lemonds
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Serial Parallel Multiplier Design in Quantum-dot Cellular Automata
abstract
An emerging nanotechnology, quantum-dot cellular automata (QCA), has the potential for attractive features such as faster speed, smaller size, and lower power consumption than transistor based technology. Quantum-dot cellular automata has a simple cell as the basic element. The cell is used as a building block to construct gates, wires, and memories. Several adder designs have been proposed, but multiplier design in QCA is a rather unexplored research area. This paper utilizes the QCA characteristics to design serial parallel multipliers. Two types of serial parallel multipliers are designed and simulated with several different operand sizes. Those designs are compared in terms of complexity, area, and latency. The serial parallel multipliers have simple and regular structures.
Heumpil Cho, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
2007 Contention-free switch-based implementation of 1024-point Radix-2 Fourier Transform Engine
abstract
This paper examines the use of a switch based architecture to implement a Radix-2 decimation in frequency fast Fourier transform engine. The architecture interconnects M processing elements with 2*M memories. An algorithm to detect and resolve memory access contention is presented. The implementation of 1024-point FFTs with 2 processing elements is discussed in detail, including timing and place-and-route results. The switch based architecture provides a factor of M speedup over a single processing element realization.
Hani Saleh, Bassam Jamil Mohd, Adnan Aziz, Earl E. Swartzlander Jr.
ICCD4
2007 The hazard-free superscalar pipeline fast fourier transform algorithm and architecture
abstract
This paper examines the superscalar pipeline Fast Fourier Transform algorithm and architecture. The algorithm presents a memory management scheme to prevent memory contention throughout the pipeline stages. The fundamental algorithm, a switch-based FFT pipeline architecture and an example 64-point FFT pipeline are presented. The proposed superscalar architecture substantially improves the FFT processing. The pipeline consists of log2N stages, where N is number of FFT points. Each stage can have M Processing Elements (PEs.) As a result, the architecture speed up is M*log2N. The pipeline algorithm is configurable to any M ≫ 1.
Bassam Jamil Mohd, Adnan Aziz, Earl E. Swartzlander Jr.
VLSI-SoC3
2006 Design of Radix-4 SRT Dividers in 65 Nanometer CMOS Technology
abstract
As technology evolves, there is a never ending need to explore design tradeoffs and alternatives. In the CMOS technologies of the recent past where minimizing the die area was crucial, radix-4 minimally redundant SRT dividers were widely used because they only require simple multiples of divisor. Quotient conversion was typically done by on-the-fly conversion. In deep submicron CMOS technology these decisions need to be reconsidered. Now it is attractive to use maximum redundancy to simplify quotient selection. Replacing the on-the-fly conversion that operates on every cycle with an adder that operates only one cycle reduces the switching factor by the order of 29x for the conversion during a double precision division. This is significant because the onthe- fly conversion can consume 30% of the total energy of a divider. Furthermore, the quotient computation is sped up by the elimination of the big lookup table of minimally redundant SRT dividers. To illustrate this concept of trading extra hardware for improved power and speed and a simpler implementation, a radix-4 maximally redundant divider is designed and implemented in 65 nm CMOS technology using an ASIC flow and single, double and triple VT devices. Clock and data gating and data recirculation techniques are used to save power. Finally, a method to evaluate design alternatives for energy efficiency is proposed that takes into account the active power consumption, the inactive power consumption and the duty cycle.
Tung N. Pham, Earl E. Swartzlander Jr.
ASAP2
2006 Systolic FFT Processors: Past, Present and Future
abstract
This paper reviews developments in the implementation of systolic fast Fourier transform processors over two decades (early 1980s to early 2000s) and identifies positive and negative lessons learned. The Modular Transform Processor was developed at TRW in 1983-84. It is a set of 6 large circuit boards that computes 4096 point FFTs using 22-bit floating-point arithmetic at sustained data rates of 40 MSPS. A single chip systolic FFT developed by the Mayo Foundation in 2001-02 computes 4096 point FFTs using 16-bit fixed-point arithmetic at sustained data rates of 200 MSPS. Some thoughts on the future directions of systolic FFT processor development are offered. Future systems will compute larger FFTs at higher data rates, will employ IEEE Standard floatingpoint arithmetic and will consume less power.
Earl E. Swartzlander Jr.
ASAP1
2005 Parallel Prefix Adder Design with Matrix Representation
abstract
The paper presents a one-shot batch process that generates a wide range of designs for a group of parallel prefix adders. The prefix adders are represented by two two-dimensional matrices and two vectors. This matrix representation makes it possible to compose two functions for gate sizing which calculate the delay and the total transistor width of the carry propagation graph of adders. After gate sizing, the critical path net-lists of the carry propagation graph are generated from the matrix representation for spice delay calculation. The process is illustrated by generating sets of delay and total transistor width pairs for 32-bit and 64-bit cases.
Youngmoon Choi, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
2005 Multiply-Accumulate Architecture for a Special Class of Optimal Extension Fields
abstract
Finite field arithmetic is useful in the implementation of error-correcting codes as well as cryptographic protocols. Large finite field numbers are particularly important in the implementation of elliptic curve cryptography. This paper presents a multiply-accumulate architecture for multipliers over a special class of type II optimal extension fields (OEFs). Type II OEFs are Galois fields GF (p/sup m/) with p a pseudo-Mersenne prime of the form p = 2/sup n/ $c, where c is "small", and an irreducible binomial of the form f (z) = z/sup m/ $2 exists over GF (p). The Type II OEF multiplier presented uses merged arithmetic to combine multiple multiply and addition operations together. Unlike previous work, the multiplier also performs subfield and extension field reduction in parallel for this class of finite fields. Though the multiplier design requires large silicon area for practical implementation, it obviates the need for performing subfield and extension field reduction separately, thereby reducing the overall delay.
Moboluwaji O. Sanu, Earl E. Swartzlander Jr.
ASAP2
2004 Parallel Montgomery Multipliers
Moboluwaji O. Sanu, Earl E. Swartzlander Jr., Craig M. Chase
ASAP2
2003 An Architecture for a Radix-4 Modular Pipeline Fast Fourier Transform
abstract
We present a radix-4 modular pipeline architecture for computing the discrete Fourier transform (DFT). For an N-point DFT, two conventional pipeline /spl radic/N-point fast Fourier transform (FFT) modules are joined by a specialized center element. The center element contains memories, coefficient ROMs, multipliers, and control logic. Compared with a standard N-point pipeline FFT, the modular FFT significantly reduces the number of delay lines to 2/spl radic/N. Further, the coefficient storage is concentrated within the center element, thereby reducing the ROM requirement within the pipeline FFT modules. The centralized memory and address generator provide data storage and reordering. The architecture has been analyzed through simulation and compared to the conventional pipeline FFT. The throughput of a standard radix-4 pipeline FFT is maintained with a slightly higher end-to-end latency. A reduction in power is achieved because the modular pipeline exhibits N/2 bit transitions on each clock as compared to y bit transitions in the conventional pipeline.
Ayman M. El-Khashab, Earl E. Swartzlander Jr.
ASAP2
2002 Implementation of a Single Chip, Pipelined, Complex, One-Dimensional Fast FourierTransform in 0.25 mu m BulkCMOS
abstract
The Mayo Foundation Special Purpose Processor Development Group (Mayo) has developed a novel fast Fourier transform (FFT) ASIC designed to operate on 16-bit complex (16-bit real, 16-bit imaginary) samples. The radix-2 FFT processor performs any power-of-two-sized transform between 2-point and 4096-point, as selected by the user. The FFT processor is wholly contained on a single 10 mm by 10 mm die implemented in 0.25 /spl mu/m bulk CMOS technology, including distributed register banks for storing all intermediate calculations, and static RAM (SRAM) for storing user programmable sine and cosine coefficients. Designed for maximum flexibility, the Mayo FFT processor includes redundant computation modules, user programmable transform length; individual, user programmable sine and cosine coefficient-storing SRAM for each of the computation modules; overflow detection and correction circuitry (in the form of user-selectable operand scaling within each computation module), 5-volt tolerant 3.3-volt I/O; and a command driven interface.
Steven M. Currie, Paul R. Schumacher, Barry K. Gilbert, Earl E. Swartzlander Jr., Barbara A. Randall
ASAP4
2002 An Analysis of the CORDIC Algorithm for Direct Digital Frequency Synthesis
abstract
The circular-mode CORDIC (coordinate rotation digital computer) algorithm is analyzed for DDFS (direct digital frequency synthesis) applications. It is shown how the CORDIC parameters should be chosen to meet given DDFS parameters. Also, three methods of CORDIC datapath quantization: rounding, truncation, and jamming, have been investigated and their error bounds are derived. Through a set of simulations, it is demonstrated that jamming has desirable characteristics in many aspects such as complexity, speed, error, and bias. Finally, it is shown that the CORDIC output can be made exact to the digits by an additional rounding process, which is especially useful for DDFS applications where the CORDIC output should be truncated to the final DAC (digital-to-analog converter) width.
Chang Yong Kang, Earl E. Swartzlander Jr.
ASAP2
2001 Analysis of Column Compression Multipliers
abstract
Column compression multipliers are frequently used in high-performance computer systems due to their short worst case delay. This paper examines the area, delay, and power characteristics of Dadda (1965) and Wallace (1964) column compression multipliers in deep submicron technology. Our analysis shows that Wallace multipliers have slightly more area and approximately the same worst case delay as Dadda multipliers. It also shows the importance of considering parasitic capacitances when determining the delay of column compression multipliers, since parasitics can increase the delay of the multiplier by over 60%. As multiplier size increases, the ratio of power to area also increases, due to longer interconnect lines and increased glitching.
K'Andrea C. Bickerstaff, Earl E. Swartzlander Jr., Michael J. Schulte
IEEE Symposium on Computer Arithmetic2
2001 A fast hybrid carry-lookahead/carry-select adder design
abstract
Article Share on A fast hybrid carry-lookahead/carry-select adder design Authors: Ohsang Kwon Sun Microsystems Inc., 901 San Antonio Rd., Pale Alto, CA Sun Microsystems Inc., 901 San Antonio Rd., Pale Alto, CAView Profile , Earl E. Swartzlander The University of Texas at Austin, Austin, Texas The University of Texas at Austin, Austin, TexasView Profile , Kevin Nowka IBM Austin Research Lab, 11400 Burnet Rd., Austin, Texas IBM Austin Research Lab, 11400 Burnet Rd., Austin, TexasView Profile Authors Info & Claims GLSVLSI '01: Proceedings of the 11th Great Lakes symposium on VLSIMarch 2001 Pages 149–152https://doi.org/10.1145/368122.368909Online:01 March 2001Publication History 9citation1,711DownloadsMetricsTotal Citations9Total Downloads1,711Last 12 Months8Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ohsang Kwon, Earl E. Swartzlander Jr., Kevin J. Nowka
ACM Great Lakes Symposium on VLSI2
2001 Time-shared TMR for fault-tolerant CORDIC processors
abstract
Presents a low-cost approach to concurrent error correction in high-performance CORDIC processors by using time-shared triple modular redundancy. Operands are partitioned into three sets of disjoint digits and operations are performed three times on different hardware components to correct possible errors by majority voting. The approach has limited latency increase and throughput reduction. Pipelining can be used to maintain the same throughput as a conventional design.
Jae-Hyuck Kwak, Vincenzo Piuri, Earl E. Swartzlander Jr.
ICASSP3
2001 DCT Implementation with Distributed Arithmetic
abstract
This paper presents an efficient method for implementing the Discrete Cosine Transform (DCT) with distributed arithmetic. While conventional approaches use the original DCT algorithm or the even-odd frequency decomposition of the DCT algorithm, the proposed architecture uses the recursive DCT algorithm and requires less area than the conventional approaches, regardless of the memory reduction techniques employed in the ROM Accumulators (RACs). An efficient architecture for implementing the scaled DCT with distributed arithmetic is also proposed. The new architecture requires even less area while keeping the same structural regularity for an easy VLSI implementation. A comparison of synthesized DCT processors shows that the proposed method reduces the hardware area of regular and scaled DCT processors by 17 percent and 23 percent, respectively, relative to a conventional design. With the row-column decomposition method, the proposed architectures can be easily extended to compute the two-dimensional DCT required in many image compression applications such as HDTV.
Sungwook Yu, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2000 A 16-Bit x 16-Bit MAC Design Using Fast 5: 2 Compressors
abstract
3:2 counters and/or 4:2 compressors have been widely used for multiplier implementations. In this paper, a new logical decomposition is derived for fast 5:2 compressor and is proposed to be used for 16-bit/spl times/16-bit MAC designs. In addition, when the accumulator output is in carry-save form, one row in partial product matrix can be eliminated prior to the partial product reduction process. These new methods are combined and explained with 16-bit/spl times/16-bit 2's complement MAC (multiply and accumulate) designs. The use of the new 5:2 compressor leads to 14% speed improvement in the MAC design over the conventional designs using 4:2 compressors and 3:2 counters.
Ohsang Kwon, Earl E. Swartzlander Jr., Kevin J. Nowka
ASAP2
2000 Fault-Tolerant Newton-Raphson and Goldschmidt Dividers Using Time Shared TMR
abstract
Iterative division algorithms based on multiplication are popular because they are fast and may utilize an already existing hardware multiplier. Two popular methods based on multiplication are Newton-Raphson and Goldschmidt's algorithm. To achieve concurrent error correction, Time Shared Triple Modular Redundancy (TSTMR) may be applied to both kinds of dividers. The hardware multiplier is divided into thirds, and the rest of the divider logic replicated around each part, to provide three independent dividers. While this reduces the size of the fault-tolerant dividers over that of traditional TMR, latency may be increased. However, both division algorithms can be modified to use lower precision multiplications during the early iterations. This saves multiply cycles and, hence, produces a faster divider.
W. Lynn Gallagher, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
2000 A Serial-Parallel Architecture for Two-Dimensional Discrete Cosine and Inverse Discrete Cosine Transforms
abstract
The Discrete Cosine and Inverse Discrete Cosine Transforms are widely used tools in many digital signal and image processing applications. The complexity of these algorithms often requires dedicated hardware support to satisfy the performance requirements of hard real-time applications. This paper presents the architecture of an efficient implementation of a two-dimensional DCT/IDCT transform processor via a serial-parallel systolic array that does not require transposition.
Hyesook Lim, Vincenzo Piuri, Earl E. Swartzlander Jr.
IEEE Trans. Computers3
2000 A Family of Variable-Precision Interval Arithmetic Processors
abstract
Traditional computer systems often suffer from roundoff error and catastrophic cancellation in floating point computations. These systems produce apparently high precision results with little or no indication of the accuracy. This paper presents hardware designs, arithmetic algorithms, and software support for a family of variable-precision, interval arithmetic processors. These processors give the programmer the ability to detect and, if desired, to correct implicit errors in finite precision numerical computations. They also provide the ability to solve problems that cannot be solved efficiently using traditional floating point computations. Execution time estimates indicate that these processors are two to three orders of magnitude faster than software packages that provide similar functionality.
Michael J. Schulte, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
1999 High-Speed CORDIC Architecture Based on Redundant Sum Formation and Overlapped s-Selection
abstract
This paper presents an architecture for accelerating CORDIC vectoring mode operations. The processing is sped up by overlapping redundant sum formation and selection of rotation direction. We analyze the latency time and area, and compare them with a conventional CORDIC implementation. The results show that the proposed scheme reduces not only the the latency but also the overall computation time. Thus, it achieves higher throughput in pipelining.
Jae Hun Choi, Jae-Hyuck Kwak, Earl E. Swartzlander Jr.
ICCD3
1999 Parallel Implementation of Multidimensional Transforms without Interprocessor Communication
abstract
Presents a modular algorithm which is suitable for computing a large class of multidimensional transforms in a general-purpose parallel environment without interprocessor communication. Since it is based on matrix-vector multiplication, it does not impose restrictions on the size of the input data as many existing algorithms do. The method is fully general, since it does not depend on the specific nature of the transform kernel and, therefore, it may be used for a wide variety of transforms. Moreover, since some 1D fast Fourier transform algorithms map the input sequence onto two or more dimensions, the new method also may be employed to efficiently compute the 1D FFT in parallel. In addition, the proposed algorithm is exploited to derive a fully systolic VLSI architecture performing multidimensional transforms, which does not need the transposer required by classical architectures.
Francescomaria Marino, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
1998 Merged Arithmetic for Computing Wavelet Transforms
abstract
A variation of merged arithmetic is applied to the implementation of the wavelet transform. This approach offers a simple design trade-off between the computational accuracy and the complexity. Our analysis shows that the trade-off is a function of the input data resolution, the number of filter taps, the arithmetic precision, and the level of the wavelet transform. The design parameter can be also fixed for a given number of taps and used to determine the minimum word size for the wavelet coefficients of the transform. The key element of this approach is to introduce a "truncation" within the merged arithmetic reduction process which provides equivalent throughput with substantially less complexity. An experiment has been conducted to verify the analysis, which suggests that 24-bit merged arithmetic is required for the EZW algorithm to handle up to a level 6-wavelet transform.
Gwangwoo Choe, Earl E. Swartzlander Jr.
Great Lakes Symposium on VLSI2
1998 A reduction scheme to optimize the Wallace multiplier
abstract
A novel bit-product reduction scheme for an n by n bit Wallace multiplier is proposed in this paper. The proposed scheme differs from the traditional Wallace method in two ways: (1) it redefines the way in which bit-products are grouped for the first stage of the bit-product reduction process, and (2) it uses a single (4,3) counter, besides the conventional half and full adders, to optimize the reduction process. The proposed method reduces the number of reduction stages when n is equal to 5, 14, 20 or 29 bits. To illustrate this new technique, the complexity and delay to reduce a 14 by 14 bit-product array using the proposed scheme are compared to that of the traditional Wallace multiplier.
Moises E. Robinson, Earl E. Swartzlander Jr.
ICCD2
1997 Power-Delay Characteristics of CMOS Multipliers
abstract
Minimizing the power consumption of circuits is important for a wide variety of applications both because of increasing levels of integration and the desire for portability. Since multipliers are widely used in computers, it is also important to maximize their speed. Frequently, the compromise between these two conflicting demands is accomplished by minimizing the product of the power dissipation and the delay. This paper reports on the dynamic power dissipation and delay of CMOS implementations of four different multipliers. Simulation was used to establish a set of models for both delay and power dissipation, and those models were then used to compute the power-delay products of the multipliers.
Thomas K. Callaway, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1997 Realization of a nonlinear digital filter on a DSP array processor
abstract
This paper presents the performance evaluation of a fast third-order Volterra digital filtering algorithm mapped onto an AT&T DSP-3 parallel processor. Five different implementations are considered. Speed-up results indicate that the "time-skewing" method is currently the fastest. An application to nonlinear communication channel equalization using a 64-QAM signal constellation is presented.
Hercule Kwan, Edward J. Powers, Earl E. Swartzlander Jr.
ASAP3
1997 Survey of low power techniques for ROMs
abstract
This paper presents a survey of low power techniques for Read Only Memories (ROMs). Significant savings in power dissipation are achieved through the use of techniques at the circuit and architecture level. The ROM circuits have been designed in 0.35 m CMOS technology and simulated using PowerMill.
Edwin de Angel, Earl E. Swartzlander Jr.
ISLPED2
1997 Hybrid CORDIC Algorithms
abstract
Each coordinate rotation digital computer iteration selects the rotation direction by analyzing the results of the previous iteration. In this paper, we introduce two arctangent radices and show that about 2/3 of the rotation directions can be derived in parallel without any error. Some architectures exploiting these strategies are proposed.
Shaoyun Wang, Vincenzo Piuri, Earl E. Swartzlander Jr.
IEEE Trans. Computers3
1996 Finite Word-Length Effects Of An Unified Systolic Array For 2-D DCT/IDCT
abstract
This paper presents a fixed-point error analysis for the unified systolic array implementation of 2-dimensional (2-D) discrete cosine transform (DCT) and 2-D inverse discrete cosine transform (IDCT). Closed form expressions for the mean and variance of fixed-point rounding-errors and truncation-errors are derived. Simulation results are provided to verify the analysis. Simulations designed to find the minimum word-length which satisfies IEEE requirements for the implementation of an 8/spl times/8 IDCT are also performed. Simulation results show that the proposed systolic array is more robust for the fixed-point error than other existing implementations for DCT/IDCT.
Hyesook Lim, Changhoon Yim, Earl E. Swartzlander Jr.
ASAP3
1996 Multidimensional systolic arrays for multidimensional DFTs
abstract
One of the most challenging problems for VLSI implementation of the discrete Fourier transform (DFT) is to efficiently implement multidimensional discrete Fourier transforms with systolic architectures. This paper presents a multidimensional systolic array for performing the multidimensional DFT. Extensions of the multidimensional systolic array are widely searched for the prime-factor computation or the 2/sup n/-point decomposed computation of one-dimensional (1-D) DFT. The essence of the proposed multidimensional systolic array is to combine different types of semi-systolic arrays into one array so that the resulting array becomes truly systolic. This systolic array does not require any preloading of input data and it produces output data at boundary PEs. No networks for intermediate spectrum transposition between constituent 1-dimensional transforms are required; therefore the entire processing is fully pipelined.
Hyesook Lim, Earl E. Swartzlander Jr.
ICASSP2
1996 Granularly-pipelined CORDIC processors for sine and cosine generators
abstract
The CORDIC algorithm is a powerful tool for computing trigonometric functions (sine and cosine) and some transcendental functions (hyperbolic sine and cosine) at a circuit complexity suited for physical implementation by using VLSI technologies. In this paper, we propose a family of architectures for high-throughput applications based on computation pipelining. The granularity of pipelining can be varied to increase the throughput at a cost of increased circuit complexity. The wide variety of solutions allows optimizing the trade-off between performance and circuit complexity, by taking into account the specific requirements and constraints of the application. The evaluations and the designer guidelines an also given.
Shaoyun Wang, Vincenzo Piuri, Earl E. Swartzlander Jr.
ICASSP3
1995 Cascaded Implementation of an Iterative Inverse--Square--Root Algorithm, with Overflow Lookahead
abstract
We present an unconventional method of computing the inverse of the square root. It implements the equivalent of two iterations of a well-known multiplicative method to obtain 24-bit mantissa accuracy. We implement each "iteration" as a separate logic module and exploit knowledge about the relative error during computation. To reduce the size of the implementation. We use overflow lookahead logic to facilitate the exponent computations. No division is required in the entire process. Examples and error analysis are given.>
Hercule Kwan, Robert Leonard Nelson Jr., Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic3
1995 Hardware Design and Arithmetic Algorithms for a Variable-Precision, Interval Arithmetic Coprocessor
abstract
This paper presents the hardware design and arithmetic algorithms for a coprocessor that performs variable-precision, interval arithmetic. The coprocessor gives the programmer the ability to specify the precision of the computation, determine the accuracy of the result, and recompute inaccurate results with higher precision. Direct hardware support and efficient algorithms for variable-precision, interval arithmetic greatly improve the speed, accuracy, and reliability of numerical computations. Performance estimates indicate that the coprocessor is 200 to 1,000 times faster than a software package for variable-precision, interval arithmetic. The coprocessor can be implemented on a single chip with a cycle time that is comparable to IEEE double-precision floating point coprocessors.>
Michael J. Schulte, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1995 Recomputing by Operand Exchanging: A Time-redundancy Approach for Fault-tolerant Neural Networks
abstract
The use of neural networks in mission-critical applications requires concurrent error detection and correction at architectural level to provide high consistency and reliability of system's outputs. Time redundancy allows for fault tolerance in digital realizations with low circuit complexity increase. In this paper, we propose the use of REcomputation with eXchanged Operands-an approach based on operands' rotation-to introduce concurrent error detection and correction, when timing constraints are not particularly strict. Different architectural approaches for neural design are considered to match the implementation constraints and to show the versatility of the proposed solutions.
Yuang-Ming Hsu, Earl E. Swartzlander Jr., Vincenzo Piuri
ASAP2
1995 A Processor for Staggered Interval Arithmetic
abstract
The paper presents the design of a high-speed processor which performs staggered interval arithmetic. Each staggered interval is represented as the sum of a set of floating point numbers plus an interval, which consists of two floating point endpoints. Staggered interval arithmetic allows the precision of the computation to be specified and the accuracy of the result to be determined. Efficient arithmetic algorithms, which reduce the number of floating point operations needed to perform staggered interval arithmetic, are introduced. To achieve high performance, the processor employs an array of pipelined floating point arithmetic units and two long accumulators. The processor provides direct hardware support for accurate and numerically reliable vector and matrix computations.
Michael J. Schulte, Earl E. Swartzlander Jr.
ASAP2
1995 An efficient systolic array for the discrete cosine transform based on prime-factor decomposition
abstract
A new design of a systolic array for computing the discrete cosine transform (DCT) based on prime-factor decomposition is presented. The basic principle of the proposed systolic array is that one-dimensional (1-D) DCT can be decomposed to a 2-dimensional (2-D) DCT by input and output index mappings and the 2-D DCT is computed efficiently on a 2-D systolic array. We modify Lee's input index mapping method in order to construct one input mapping table instead of three input index mapping tables. The proposed systolic array avoids the need for the array transposer that was required by earlier implementations for the prime-factor DCT algorithms, and thus all processing can be pipelined. The proposed design of systolic array provides a simple and regular structure, which is well suited for VLSI implementation.
Hyesook Lim, Earl E. Swartzlander Jr.
ICCD2
1995 A coprocessor for accurate and reliable numerical computations
abstract
This paper presents the architecture and hardware design of a special-purpose coprocessor that performs variable-precision, interval arithmetic. Variable-precision arithmetic allows the precision of the computation to be specified, based on the problem to be solved and the required accuracy of the results. Interval arithmetic produces two values for each result, such that the true result is guaranteed to be between the two values. The coprocessor gives the programmer the ability to specify the precision of the computation, determine the accuracy of the results, and recompute inaccurate results with higher precision. Direct hardware support for variable-precision, interval arithmetic greatly improves the accuracy and reliability of numerical computations. Execution time estimates indicate that the coprocessor is two to three orders of magnitude faster than an existing software package for variable-precision, interval arithmetic.
Michael J. Schulte, Earl E. Swartzlander Jr.
ICCD2
1995 Time-Redundant Multiple Computation for Fault-Tolerant Digital Neural Networks
abstract
In mission-critical applications of artificial neural networks, error correction at the architectural level is often mandatory to guarantee consistency and reliability of the network's outputs. Time redundancy allows for fault tolerance with low circuit complexity overhead. In this paper, the application REcomputing with Triplication With Voting (RETWV) at the system level is proposed for concurrent error correction in neural networks. Feed-forward multi-layered neural networks are considered as an example, but the proposed technique can be easily extended to different neural paradigms.
Yuang-Ming Hsu, Vincenzo Piuri, Earl E. Swartzlander Jr.
ISCAS3
1995 Fault-Tolerant Neural Architectures: The Use of Rotated Operands
Yuang-Ming Hsu, Vincenzo Piuri, Earl E. Swartzlander Jr.
ISCAS3
1995 Merged CORDIC Algorithm
abstract
The COordinate Rotation DIgital Computer (CORDIC) algorithm is an iterative procedure to evaluate various elementary functions. It usually consists of one scaling multiplication and n+1 elementary shift-add iterations in an n bit processor. These iterations can be paired off to form double iterations to lower the hardware complexity while the computational complexity stays the same. With this structure, the shifter size is reduced to 1/2 (1+9/n+1). In this paper, we present this merged algorithm, its error analysis, and software simulation results.
Shaoyun Wang, Earl E. Swartzlander Jr.
ISCAS2
1995 Rapid prototyping fault-tolerant heterogeneous digital signal processing systems
abstract
An approach is presented that permits the configuration of application specific hardware, with arbitrary hardware redundancy, to match the signal flow graph of arbitrary applications. The hardware is mapped to the signal flow graph of an application. An inventory of heterogeneous processors, specialized to perform a predefined set of functions, enables rapid prototyping of systems with arbitrary topologies. Application specific systems that match the signal flow graph of applications outperform general purpose systems in speed and throughput. This research focuses on solving the problems associated with the interconnection of the heterogeneous building blocks. A communication architecture is proposed that allows the interconnection of processors with varying speed and functionalities. Standardization of the interface control unit (ICU) greatly reduces the development cost by removing the need to design custom interfaces. The ICU permits the introduction of varying degree of hardware redundancy into the topology at the system, cluster or processor level.
Mohammad S. Khan, Earl E. Swartzlander Jr.
RSP2
1994 A systolic array for 2-D DFT and 2-D DCT
abstract
A new approach for computing the 2-D DFT (discrete Fourier transform) and 2-D DCT (discrete cosine transform) is presented. A new design of a systolic array for transposed matrix multiplication is also shown in this paper. The new 2-D DFT/DCT avoids the need for the array transposer that was required by earlier implementations, and all processing can be pipelined easily. This approach employs a simple and regular structure that is well suited for VLSI implementation. This array can be easily scaled without modifying the basic control scheme and PE structure.>
Hyesook Lim, Earl E. Swartzlander Jr.
ASAP2
1994 A variable-precision interval arithmetic processor
abstract
This paper presents a special-purpose processor which implements variable-precision, interval arithmetic. Variable-precision arithmetic allows the precision of the computation to be specified, based on the problem to be solved and the required accuracy of the computation. Interval arithmetic produces two values for each result, such that the true result is guaranteed to be between the two values. The distance between the two values gives an upper bound on the error. Direct hardware support for variable-precision, interval arithmetic greatly improves the accuracy of the computation, and is much faster than existing software methods for controlling numerical error. Area and delay estimates indicate that the processor can be implemented on a single chip with a cycle time which is comparable to existing IEEE double-precision floating point processors. For computationally intensive problems, an application-specific array of variable-precision, interval arithmetic processors can execute in parallel to provide high-performance and numerically reliable results.>
Michael J. Schulte, Earl E. Swartzlander Jr.
ASAP2
1994 A New Asynchronous Multiplier Using Enable/Disable CMOS Differential Logic
abstract
This paper presents a technique for asynchronous logic design using ECDL (Enable/Disable CMOS Differential Logic). A pipelined serial-parallel multiplier clocked at 55.6 MHz has been designed to show the implementation of this technique. The serial-parallel multiplier architecture has been designed in ECDL using MAGIC, and circuit simulations have been done in HSPICE using a 2 /spl mu/m model from MOSIS. An evaluation of the area using ECDL is presented and compared against techniques used in the past to show that a significant reduction in area overhead is possible.>
Edwin de Angel, Earl E. Swartzlander Jr., Jacob A. Abraham
ICCD2
1994 What Types of Research Papers Should We Be Writing?
Thomas L. Casavant, Chi-Yuan Chin, Wen-Tsuen Chen, Kang G. Shin, Earl E. Swartzlander Jr., Joseph E. Urban
ICPADS5
1994 Is It Possible to Fairly Compare Interconnection Networks?
José Duato, C. T. Howard Ho, Ferng-Ching Lin, Lionel M. Ni, Earl E. Swartzlander Jr.
ICPADS5
1994 Sorting Networks with Built-In Error Correction
abstract
A sorting network with built-in error correction is proposed in this paper. A time shared TMR scheme is used to achieve the error correcting capability. A quarter of the original sorting network based on perfect shuffle is triplicated and voted in each stage. The hardware complexity of this time shared TMR error correcting sorting network is a little more than the original sorting network. The price is that the delay time increases by a factor of 4. However, the throughput penalty can be minimized by pipelining. A technology-independent gate level analysis of hardware complexity and delay time is included in this paper. Possible variations of the basic design are also discussed.
Yuang-Ming Hsu, Earl E. Swartzlander Jr.
ICPADS2
1994 Heterogeneous Parallel Computing
abstract
Heterogeneous parallel computing (i.e., parallel computers where the processors are customized or tailored to efficiently execute specific.classes of algorithms) is the only practical way to solve many computationally intensive problems. In contrast to array computers constructed with identical general purpose processors, heterogeneous parallel computing achieves very high levels of throughput, small size, and low power. The improvement in the speed-size product is often in excess of two orders of magnitude. This talk reviews past endeavors in application specific processing (which can be viewed as the root of heterogeneous computing), current research in heterogeneous parallel computing, and offers the prediction that in the near future “system compilation” (analogous to “silicon compilation”) will greatly facilitate the development of heterogeneous parallel computers.
Earl E. Swartzlander Jr.
ICPADS1
1994 A Standardized Interface Control Unit for Heterogeneous Digital Signal Processors
abstract
An approach is presented for the implementation of high-performance heterogeneous digital signal processors. This approach facilitates tailoring the configuration of the hardware to the signal flow graph of the application. The hardware system comprises heterogeneous processors which are specialized for performing a predefined set of functions. The focus of this research is to solve the communication problems associated with the interconnection of the heterogeneous building blocks to implement application-specific systems. An interface control unit (ICU) architecture is proposed that allows the interconnection of processors with varying speeds and functionalities. A protocol has been developed that facilitates transfer of large data blocks in sparse networks. The hardware design and implementation of the ICU is based on this protocol. The digital design of the ICU is suitable for fabrication as a single VLSI circuit.>
Mohammad S. Khan, Earl E. Swartzlander Jr.
ISCAS2
1994 Boundary scan in board manufacturing
Thomas A. Ziaja, Earl E. Swartzlander Jr.
J. Electron. Test.2
1994 Hardware Designs for Exactly Rounded Elemantary Functions
abstract
This paper presents hardware designs that produce exactly rounded results for the functions of reciprocal, square-root, 2/sup x/, and log/sub 2/(x). These designs use polynomial approximation in which the terms in the approximation are generated in parallel, and then summed by using a multi-operand adder. To reduce the number of terms in the approximation, the input interval is partitioned into subintervals of equal size, and different coefficients are used for each subinterval. The coefficients used in the approximation are initially determined based on the Chebyshev series approximation. They are then adjusted to obtain exactly rounded results for all inputs. Hardware designs are presented, and delay and area comparisons are made based on the degree of the approximating polynomial and the accuracy of the final result. For single-precision floating point numbers, a design that produces exactly rounded results for all four functions has an estimated delay of 80 ns and a total chip area of 98 mm/sup 2/ in a 1.0-micron CMOS technology. Allowing the results to have a maximum error of one unit in the last place reduces the computational delay by 5% to 30% and the area requirements by 33% to 77%.>
Michael J. Schulte, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
1993 Estimating the power consumption of CMOS adders
abstract
Six types of adders are examined in an attempt to model their power dissipation. It is shown that the use of a relatively simple model provides results that are qualitatively accurate, when compared to more sophisticated models and to physical implementations of the circuits. The main discrepancy between the simple model and the physical measurements seems to be the assumption that all gates will consume the same amount of power when they switch, regardless of their fan-in or fanout. Because the carry lookahead adder has several gates with a fan out and fan-in higher than two, the simple model underestimates its power dissipation.>
Thomas K. Callaway, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1993 Exact rounding of certain elementary functions
abstract
An algorithm is described which produces exactly rounded results for the functions of reciprocal, square root, 2/sup x/, and log 2/sup x/. Hardware designs based on this algorithm are presented for floating point numbers with 16- and 24-b significands. These designs use a polynomial approximation in which coefficients are originally selected based on the Chebyshev series approximation and are then adjusted to ensure exactly rounded results for all inputs. To reduce the number of terms in the approximation, the input interval is divided into subintervals of equal size and different coefficients are used for each subinterval. For floating point numbers with 16-b significands, the exactly rounded value of the function can be computed in 51 ns on a 20-mm/sup 2/ chip. For floating point numbers with 24-b significands, the functions can be computed in 80 ns on a 98-mm/sup 2/ chip.>
Michael J. Schulte, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1993 Reduced area multipliers
abstract
As developed by Wallace (1964) and Dadda (1965), a high-speed method for the parallel multiplication of two binary numbers is to reduce their partial products to two numbers whose sum is equal to the product. The resulting two numbers are then summed using a fast carry-propagate adder. The authors present a multiplier, the reduced area multiplier, with a novel reduction scheme which results in fewer components and less interconnect overhead than either Wallace or Dadda multipliers. This reduction scheme is especially useful for pipelined multipliers, because it minimizes the number of latches required in the reduction of the partial products. Equations are given for determining the number of components and a method is presented for estimating the interconnect overhead for Wallace, Dadda and reduced area multipliers. Area estimates indicate that pipelined reduced area multipliers require 3 to 8% less area than equivalent Wallace multipliers and 15 to 25% less area than equivalent Dadda multipliers.>
K'Andrea C. Bickerstaff, Michael J. Schulte, Earl E. Swartzlander Jr.
ASAP3
1993 A Comparative Evaluation of Adders Based on Performance and Testability
abstract
Testability is becoming an increasing concern in the design of present-day VLSI systems because of their higher density and complexity. This is particularly important in the case of arithmetic units, such as adders, which form the core of any processing unit. Techniques like design for testability (DFT) have been implemented, but a methodology for evaluating and selecting a suitable adder has not been developed. We present an exhaustive comparison of adders in terms of performance, area and testability, by formulating a figure of merit, the PLUS factor. The results of this comparison can be extended to evaluate the suitability of an adder for a particular set of design goals and constraints.>
Rathish Jayabharathi, Thomas Thomas, Earl E. Swartzlander Jr.
ICCD3
1993 Superpipelined Adder Designs
Ishaq H. Unwala, Earl E. Swartzlander Jr.
ISCAS2
1993 Design and implementation of an interface control unit for rapid prototyping
abstract
A major difficulty in rapid prototyping is the interconnection of processors with tailored networks to implement systems. This difficulty may be alleviated by utilizing a standardized processor-to-processor interface. This paper describes an interface control unit (ICU) VLSI circuit which uses a 'standard' communication interface to implement signal processing systems. The ICU comprises queuing structures, communication channel interfaces, a processor interface, and a control interface. Communication protocols are described for interfaces between the ICU and communication channels, the processor, and the control unit. Based on these protocols, a digital design has been completed which is described. The chip will be fabricated using 0.6 micron CMOS technology.>
Mohammad S. Khan, Earl E. Swartzlander Jr.
RSP2
1993 Modified Booth algorithm for high radix fixed-point multiplication
abstract
It is shown that when the standard Booth multiplication algorithm is extended to higher radix (>2) fixed-point multiplication, incorrect results are produced for some word sizes. A rule which modifies the algorithm to correct this problem is presented. The modification is defined for multipliers of any size, with any power of two radix.>
Philip E. Madrid, Brian Millar, Earl E. Swartzlander Jr.
IEEE Trans. Very Large Scale Integr. Syst.3
1992 Advanced technology for improved signal processor efficiency
abstract
Wafer scale integration technology offers the promise of implementing application specific processors with significantly higher data rates, lower power, and smaller size than conventional VLSI implementations. Wafer scale integration implementations replace most of the signal lines between chips with intra-wafer lines that exhibit one to two orders of magnitude less stray capacitance so they may be driven at higher rates while consuming much less power. Application specific processors implemented with regular arrays of processing elements are attractive because their regularity simplifies the design, fabrication, and circumvention of faulty elements. This paper shows that one dimensional systolic arrays are more attractive for this application than other regular architectures. This paper also shows that (1:N) and (M:N) pooled sparing at the macrocell level is feasible to overcome the defects implicit in the fabrication process. Finally an example design for a systolic FFT processor is described to illustrate the wafer scale implementation of a signal processor.>
Earl E. Swartzlander Jr.
ASAP1
1992 Arithmetic Error Analysis of a new Reciprocal Cell
abstract
The arithmetic error of a fast reciprocal 16-b cell is analyzed. This VLSI cell computes the result in two clock cycles by the use of a very small ROM table and innovative second-order interpolation. The reliability of the predictive formulas (for the arithmetic errors of the cell) is demonstrated by comparing the predictions with computer simulation results. The methodology outlined can easily be extended to other VLSI cells.>
Vijay K. Jain, Gibert E. Perez, Earl E. Swartzlander Jr.
ICCD3
1992 Modified Booth Algorithm for High Radix Multiplication
abstract
It is shown that, in general, the standard Booth algorithm cannot be extended to higher radix (>2) multiplication. A rule to modify the Booth standard radix-2 algorithm for higher-radix multiplication is presented. This rule corrects the product computed by Booth's algorithm for certain cases of high-radix bit-recoded multiplications. In addition, the modification is defined for multipliers of any size, utilizing any power-of-2-bit recoding.>
Philip E. Madrid, Brian Millar, Earl E. Swartzlander Jr.
ICCD3
1992 A Spanning Tree Carry Lookahead Adder
abstract
The design of the 56-b significant adder used in the Advanced Micro Devices Am29050 microprocessor is described. Originally implemented in a 1- mu m design role CMOS process, it evaluates 56-b sums in well under 4 ns. The adder employs a novel method for combining carries which does not require the back propagation associated with carry lookahead, and is not limited to radix-2 trees, as is the binary lookahead carry tree of R.P. Brent and H.T. Kung (1982). The adder also utilizes a hybrid carry lookahead-carry select structure which reduces the number of carriers that need to be derived in the carry lookahead tree. This approach produces a circuit well suited for CMOS implementation because of its balanced load distribution and regular layout.>
Thomas W. Lynch, Earl E. Swartzlander Jr.
IEEE Trans. Computers2
1991 The redundant cell adder
abstract
The design of the 56-b significand adder for the Advanced Micro Devices, Am29050 microprocessor, is described. This is a 1- mu m design rule CMOS realization of a high-performance RISC (reduced instruction set computer) microprocessor that implements IEEE Standard 754 floating-point arithmetic. To achieve an add time of under 4 ns for the 56-b significand and to avoid multistage pipelines which significantly impair compiler efficiency, a redundant cell adder has been developed. This redundant cell adder design combines carry lookahead adders realized with Manchester carry chains and the carry select adder concept to achieve approximately twice the speed of the traditional carry lookahead adder. This adder achieves a 3.2-ns measured add time for 56-bit operands and is of reasonable size.>
Thomas W. Lynch, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic2
1991 High-speed multiplier design using multi-input counter and compressor circuits
abstract
The design of a fast multiplier implemented using either
Mayur Mehta, Vijay Parmar, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic3
1991 Arithmetic for digital neural networks
abstract
The implementation of large input digital neurons using designs based on parallel counters is described. The implementation of the design uses a two-cell library, in which each cell is implemented using switching trees which are pipelined binary trees of n-channel transistors. Results obtained from initial switching trees realized with a 3- mu m CMOS process indicate that the design is capable of being pipelined at 40 MHz sample rates, with better performance expected for more advanced technologies. It appears feasible to develop a wafer-scale implementation with 2000 neurons (each with 1000 inputs) that would perform 3*10/sup 12/ additions/s.>
David Zhang 0001, Graham A. Jullien, William C. Miller, Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic4
1991 The case for application specific computing
abstract
Application specific computing is the only way to solve many computationally intensive problems. In contrast to general purpose computing, application specific computing can achieve high throughput, small size, and (for CMOS realizations) low power. The improvement in the area time product is often in excess of two orders of magnitude. This paper reviews past endeavors in special purpose processing, current research in application specific array processing, and offers the prediction that in the near future 'system compilation' will greatly facilitate the development of application specific computers.>
Earl E. Swartzlander Jr.
ASAP1
1985 Arithmetic for high speed FFT implementation
abstract
This paper describes recent progress in the implementation of high speed spectrum analysis systems with state-of-the-art commercial and semi-custom VLSI circuits. Initial efforts are producing Fast Fourier Transform (FFT) and inverse FFT processors that operate at data rates of up to 40 MHz (complex). The current implementation computes transforms of up to 16,384 points in length by means of the radix 4 pipeline FFT algorithm. The interstage reordering is performed by delay commutators implemented with semi-custom VLSI, while the arithmetic is performed by commercial single chip 22 bit floating point adders and multipliers. This paper explains the pipeline FFT implementation and focuses attention on the arithmetic used to realize the design.
Earl E. Swartzlander Jr., John A. Eldon
IEEE Symposium on Computer Arithmetic1
1985 Image processing address generator chip
abstract
Accurate high speed rotation, warpage, translation, or rescaling of a two-dimensional image requires large RAMs, fast multiplier accumulators (MACs), and sophisticated address generators and controllers. TRW is designing a CM36 integrated circuit that generates the necessary control signals and data and coefficient addresses, economically replacing roughly 100 MSI and SSI components. The chip's target speed of 10 MHz is well matched to commercially available memories and MACs. With its versatile instruction set, the chip efficiently supports all first and second order image transforms, plus two dimensional filtering with a kernel size of up to 225 pixels.
John A. Eldon, Zoltan Stroll, Earl E. Swartzlander Jr.
ICASSP3
1985 Foreword: Advances in Distributed Computing Systems
Stephen F. Lundstrom, Earl E. Swartzlander Jr.
IEEE Trans. Software Eng.2
1984 Fast transform processor implementation
abstract
This paper describes recent progress in implementation of a 40 MHz (complex) data rate frequency domain adaptive digital filter. The filter uses multiple time overlapped channels each consisting of an FFT, a frequency domain multiplier, and an inverse FFT. The 4096 point FFT and inverse FFT processors realize the McClellan and Purdy radix 4 pipeline FFT algorithm with 22 bit floating point arithmetic. The arithmetic is performed with single chip floating point adders and multipliers. The interstage reordering is performed with a delay commutator implemented with semi-custom VLSI. By using state of the art arithmetic components and judicious semi-custom circuit development, an FFT processor has been implemented that computes a 4096 point (complex) transform in 102 microseconds.
Earl E. Swartzlander Jr., George Hallnor
ICASSP1
1983 Digital signal processing with VLSI technology
abstract
During the last decade, a new generation of integrated circuits has been developed that is directly applicable to the implementation of advanced signal processors. Examples of such circuits include microprocessors, fast wide word memories, single chip multipliers, floating point adders, etc. Although important in their own right, as examples of advanced technology, such circuits are most significant as components for the development of more complex structures. This paper shows how one such structure, a high performance digital filter, is implemented using current technology. Specifically, a processor that performs on the order of one billion radix 2 butterflies per second is shown to be feasible. Such high levels of performance are required to realize advanced digital signal processing systems such as adaptive beam formers. Future VLSI device research should be guided by the experience gained in the course of designs such as this.
Earl E. Swartzlander Jr., Louis S. Lome, George Hallnor
ICASSP1
1983 Sign/Logarithm Arithmetic for FFT Implementation
abstract
Sign/logarithm arithmetic is applicable to a variety of numerical applications where wide dynamic range and small wordsize are required. In this paper the basic sign/logarithm arithmetic operations required for signal processing (i.e., addition, subtraction, and multiplication) are reviewed, the computational errors are analyzed for FFT realization, and simulation results are presented which serve to verify the analysis. It is shown that the sign/logarithm approach provides improved arithmetic quantization error performance for a given word size over FFT's implemented with conventional fixed or floating point arithmetic, and that the sign/logarithm implementation is faster and less complex than conventional approaches.
Earl E. Swartzlander Jr., D. V. Satish Chandra, H. Troy Nagle, Scott A. Starks
IEEE Trans. Computers1
1982 Supersystems: Technology and Architecture
abstract
Supersystems are general purpose computers achieving throughputs in excess of 1 billion instructions per second (BIPS). As such, supersystems present formidable design challenges in the areas of technology and architecture. This paper examines three of the classical design options: high-speed monoprocessors, array processors, and distributed processors. The latter approach appears most desirable for supersystems, but will require improved interconnection networks.
Earl E. Swartzlander Jr., Barry K. Gilbert
IEEE Trans. Computers1
1980 Signal processing architectures with VLSI
abstract
VLSI technology currently allows construction of integrated circuits with thousands of gates of operating at clock rates approaching 100 MHz. Effective use of VLSI requires careful coordination of the system and device architecture. Architectural considerations are discussed and recent VLSI chips are examined to illustrate the issues.
Earl E. Swartzlander Jr.
ICASSP1
1980 Merged Arithmetic
abstract
The concept of merged arithmetic is introduced and demonstrated in the context of multiterm multiplication/addition. The merged approach involves synthesizing a composite arithmetic function (such as an inner product) directly instead of decomposing the function into discrete multiplication and addition operations. This approach provides equivalent arithmetic throughput with lower implementation complexity than conventional fast multipliers and carry look-ahead adder trees.
Earl E. Swartzlander Jr.
IEEE Trans. Computers1
1980 Arithmetic for Ultra-High-Speed Tomography
abstract
The first of a new generation of high performance X-ray computed tomographic (CT) machines, the Dynamic Spatial Reconstructor, imposes a requirement for digital signal processing rates which are 3–4 orders of magnitude greater than the capability of current X-ray computed tomography processors. To solve the large-scale computational problems for this and similar CT units which are currently under development, three candidate arithmetic implementations of ultra-high-speed convolutional filtering and weighted linear summation algorithms have been developed and compared. Since both convolution and weighted summation are performed via the inner product operation, which is the basis for most digital signal processing algorithms, the results are widely applicable. The three arithmetic approaches are a two's complement modular array, a merged arithmetic module, and a sign/logarithm convolver. A figure of merit, which relates processing speed to complexity, is used to compare the three arithmetic approaches. It is demonstrated that processing rates in the billions of multiply-add operations per second may be realized with special-purpose processors of moderate complexity.
Earl E. Swartzlander Jr., Barry K. Gilbert
IEEE Trans. Computers1
1979 Microprogrammed Control for Specialized Processors
abstract
Several microprogramming techniques suitable for the control of specialized processors are presented. The controllers are developed by restricting the next state function of a classical Moore machine, and by using microprogramming in their implementation The concept is well matched to applications which do not require complex branching logic, and is optimized for efficient implementation with commercially available integrated circuits. These are demonstrated by an example controller which illustrates the concept.
Earl E. Swartzlander Jr.
IEEE Trans. Computers1
1979 Comment on "The Focus Number System"
abstract
In a recent correspondence1Lee and Edgar present a logarithmic number system and describe algorithms for the four basic arithmetic operations. Although it is encouraging to see continued interest in the area of specialized number systems, Lee and Edgar's work represents a duplication of work published in this TRANSACTIONS in 1975[1].
Earl E. Swartzlander Jr.
IEEE Trans. Computers1
1979 A Routing Algorithm for Signal Processing Networks
abstract
Algorithms are described here which are suitable for control of a distributed network of signal processors which are interconnected with dedicated paths using decentralized routing control. This paper discusses the algorithm used to implement the processing required at the switching node processor, which can be realized with any of several LSI technologies. Although centralized systems are more efficient in terms of hardware, a single failure in the controller may disable the entire network. This distributed network is implemented with a crossbar switch at each node which facilitates the simultaneous utilization of multiple paths in the network. Index Terms-Computer networks, data routing algorithms, digital signal processing, distributed processing, signal processing networks.
Earl E. Swartzlander Jr., Douglas J. Heath
IEEE Trans. Computers1
1978 Merged arithmetic for signal processing
abstract
The concept of merged arithmetic is introduced and applied to signal processing. The basic idea involves synthesizing a composite arithmetic function (e.g., a complex multiply) directly instead of decomposing it into multiply and add operations as is conventional practice. This approach results in a simpler design which is also faster.
Earl E. Swartzlander Jr.
IEEE Symposium on Computer Arithmetic1
1978 Inner Product Computers
abstract
The inner product computer is a special-purpose computational unit intended to be used as an adjunct to a general-purpose digital computer to perform numerical processing tasks which previously exceeded the capacity of the general-purpose computer. The algorithmic structure of the inner product is briefly reviewed in the first section of this paper. Methods are described for computing the inner product of complex vectors with a series of four real inner products. Several hardware implementations of the inner product computer are described and then compared in terms of speed and complexity; a figure of merit is developed to simplify the comparison. The utility of this computational unit is demonstrated via the examination of a large-scale numerical problem, computerized three-dimensional x-ray reconstruction (computerized tomography) arising in the biomedical sciences. Finally, a comparison is given of the size of a general-purpose computer required to execute a large-scale processing task with that of an inner product computer to execute the same task. The inner product computer greatly reduces computational costs for the solution of a large class of problems, including computerized tomography, image restoration, weather forecasting, and economic modeling.
Earl E. Swartzlander Jr., Barry K. Gilbert, Irving S. Reed
IEEE Trans. Computers1
1975 The Sign/Logarithm Number System
abstract
A signed logarithmic number system, which is capable of representing negative as well as positive numbers is described. A number is represented in the sign/logarithm number system by a sign bit and the logarithm of the absolute value of the number (scaled to avoid negative logarithms).
Earl E. Swartzlander Jr., Aristides G. Alexopoulos
IEEE Trans. Computers1
1973 The Quasi-Serial Multiplier
abstract
A novel technique for digital multiplication is presented that represents a considerable departure from conventional (i.e., add and shift or fully parallel) multiplication algorithms. The quasi-serial multiplier generates the bits of the product sequentially from least significant to most significant. Each bit is computed by "counting" the number of ones in the corresponding column of the bit-product matrix and adding the previous carrys. This single operation yields both the product bit and the carrys for the next column. The quasi-serial multiplier requires 2n of these count and add operations to determine the product of two n-bit numbers.
Earl E. Swartzlander Jr.
IEEE Trans. Computers1
1973 Parallel Counters
abstract
Multiple-input circuits that count the number of their inputs that are in a given state (normally logic ONE) are called parallel counters. In this paper three separate types of counters are described, analyzed, and compared. The first counter consists of a network of full adders. The second counter uses a combination of full adders and fast adders (that may be realized with READ-ONLY memories), while the third type of counter uses quasi-digital (i.e., analog current summing) techniques to generate an analog signal proportional to the count which is then digitized.
Earl E. Swartzlander Jr.
IEEE Trans. Computers1
1973 The inner product computer (Ph.D. Thesis abstr.)
Earl E. Swartzlander Jr.
IEEE Trans. Inf. Theory1
1973 Review of "Introduction to Mathematical Techniques in Pattern Recognition" by Harry C. Andrews
Earl E. Swartzlander Jr.
IEEE Trans. Syst. Man Cybern.1
1973 Review of "Fundamentals of Pattern Recognition" by Edward A. Patrick
Earl E. Swartzlander Jr.
IEEE Trans. Syst. Man Cybern.1