EDBT 2026 Demo / reviewers in the wild / expert
Elisardo Antelo
dblp:48/2572
· DBLP profile ↗
38ranked-venue papers
14as first author
0since 2021 · last 2019
0000-0003-3743-3689ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 10 first-authorTheory of computation · 12 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Integrated circuit design · 64% Processor architecture and microarchitecture · 23% Electronic design automation · 5% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.4 | 3 | 2014 | Fast Radix-10 Multiplication Using Redundant BCD Codes · IEEE Trans. Computers 2014 Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012 Improved Design of High-Performance Parallel Decimal Multipliers · IEEE Trans. Computers 2010 |
Integrated circuit design › digital arithmetic circuits
CORDIC |
0.3 | 7 | 2012 | Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012 A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008 Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 |
Integrated circuit design › digital arithmetic circuits › parallel multiplication
partial product reduction |
0.3 | 2 | 2014 | Fast Radix-10 Multiplication Using Redundant BCD Codes · IEEE Trans. Computers 2014 Reducing the Computation Time in (Short Bit-Width) Two's Complement Multipliers · IEEE Trans. Computers 2011 |
Integrated circuit design
digital circuit design |
0.2 | 8 | 2011 | Reducing the Computation Time in (Short Bit-Width) Two's Complement Multipliers · IEEE Trans. Computers 2011 A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008 Error Analysis and Reduction for Angle Calculation Using the CORDIC Algorithm · IEEE Trans. Computers 1997 |
Processor architecture and microarchitecture
computer arithmetic |
0.2 | 5 | 2008 | A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008 Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 |
Electronic design automation › hardware security
hardware signatures |
0.1 | 1 | 2012 | FlexSig: Implementing flexible hardware signatures · ACM Trans. Archit. Code Optim. 2012 |
Integrated circuit design › digital arithmetic circuits
special function unit |
0.1 | 1 | 2012 | Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2012 | FlexSig: Implementing flexible hardware signatures · ACM Trans. Archit. Code Optim. 2012 |
Integrated circuit design › digital arithmetic circuits › integer multiplier
two's complement multiplier |
0.1 | 1 | 2011 | Reducing the Computation Time in (Short Bit-Width) Two's Complement Multipliers · IEEE Trans. Computers 2011 |
Integrated circuit design
digital arithmetic circuits |
0.1 | 2 | 2003 | Radix-4 Reciprocal Square-Root and Its Combination with Division and Square Root · IEEE Trans. Computers 2003 CORDIC Vectoring with Arbitrary Target Value · IEEE Trans. Computers 1998 |
Processor architecture and microarchitecture › arithmetic unit
arithmetic unit design |
0.1 | 5 | 2008 | A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008 Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 |
Processor architecture and microarchitecture › division algorithm
digit-recurrence divider |
0.1 | 1 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture
division algorithm |
0.1 | 1 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
GPUs and heterogeneous computing
graphics accelerator |
0.1 | 1 | 2005 | High-Throughput CORDIC-Based Geometry Operations for 3D Computer Graphics · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › division algorithm
quotient digit selection |
0.1 | 1 | 2005 | Digit-Recurrence Dividers with Reduced Logical Depth · IEEE Trans. Computers 2005 |
Processor architecture and microarchitecture › computer arithmetic
digit-recurrence algorithm |
0.0 | 2 | 2000 | Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/Vectoring · IEEE Trans. Computers 2000 Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with Scaling · IEEE Trans. Computers 1998 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2012 | FlexSig: Implementing flexible hardware signatures · ACM Trans. Archit. Code Optim. 2012 |
Integrated circuit design › digital arithmetic circuits
division and square root |
0.0 | 1 | 2003 | Radix-4 Reciprocal Square-Root and Its Combination with Division and Square Root · IEEE Trans. Computers 2003 |
Processor architecture and microarchitecture › microprocessor design › processor core design
embedded cores |
0.0 | 1 | 2011 | Reducing the Computation Time in (Short Bit-Width) Two's Complement Multipliers · IEEE Trans. Computers 2011 |
Processor architecture and microarchitecture
arithmetic unit |
0.0 | 2 | 1997 | Error Analysis and Reduction for Angle Calculation Using the CORDIC Algorithm · IEEE Trans. Computers 1997 Unified Mixed Radix 2-4 Redundant CORDIC Processor · IEEE Trans. Computers 1996 |
Integrated circuit design › digital circuit design › arithmetic circuit design
division and square root unit |
0.0 | 1 | 1998 | Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with Scaling · IEEE Trans. Computers 1998 |
Processor architecture and microarchitecture › pipelining
pipelined processor |
0.0 | 1 | 1996 | Unified Mixed Radix 2-4 Redundant CORDIC Processor · IEEE Trans. Computers 1996 |
Methods — techniques the papers use, named apart from their topics
carry-save addition · 0.3signed-digit radix-10 recoding · 0.2sign estimation · 0.1dynamic resource redistribution · 0.1bloom filter · 0.1radix-4 modified booth encoding · 0.1logic synthesis · 0.1scale factor compensation · 0.1signed-digit recoding · 0.1multioperand carry-save addition · 0.1CORDIC · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | New 3D Projection Transformation for Point Cloudsabstract3D Computer Graphics based on clouds of points is an alternative of interest for high quality rendering of complex scenes. For high quality graphics, the computational requirements are significantly higher than in the case of conventional front-end vertex processing. One of the key computations in the graphics pipeline is the projection transformation. Conventional computation of the point-based projection using current vertex processors might not be the best option for scalable, high quality 3D graphics based on point clouds. In this work we propose a new scalable method that takes advantage of the special characteristics of the point rendering model. The number of reciprocal/division computations for perspective correction is reduced, using instead more effective multiply-accumulate operations. This is done by performing a linear approximation of the reciprocal at the cost of introducing an error in the pixel value. There is a trade-off in the error introduced and the exactness of the computation (linear approximation instead of reciprocal). Álvaro Vázquez, Elisardo Antelo |
ARITH | 2 |
| 2017 | A Number System Approach for Adder TopologiesabstractThe design space exploration for fast and power efficient adders is of increasing interest for microprocessors and graphic and digital signal processors. Recently, several methods have been proposed to explore the design space of adders, where well known designs appear as possible instances. These methods are based on the identification of parameters that lead to different hardware structures. In this work, we go a step further by exploring the mathematical foundation behind the trees for carry computation. We propose an algorithm that allows to obtain any adder topology based on design decisions. The method is based on finding representations of integers in a given number system. This leads to an adder model that allows the design of any adder structure in a compact a formal way. The proposed formal model might be useful for a formal design description of adders and it can be incorporated to CAD tools. Álvaro Vázquez, Elisardo Antelo |
ARITH | 2 |
| 2017 | A Sum Error Detection Scheme for Decimal ArithmeticabstractUsers of financial and e-commerce services demand a high degree of reliability and at the same time an increasing demand of speed of processing. On the other hand, soft errors are becoming more significant due to the higher densities and reduced CMOS integration technology sizes. Among the basic arithmetic operations, addition/subtraction is the most demanded. Although in the past, binary implementations were only considered, today decimal implementations are becoming important. In this context, we introduce a modular design for fast error checking of binary and decimal (BCD) addition/subtraction operations that avoids the whole replication of the arithmetic units. Unlike other error checkers based on parity prediction or residue checking, this is a separable design that lies completely off of the critical path of the protected adder without incurring in important penalties in area or performance. Álvaro Vázquez, Elisardo Antelo |
ARITH | 2 |
| 2016 | Asymmetric Allocation in a Shared Flexible Signature Module for Multicore ProcessorsabstractHardware signatures based on Bloom filters are used to support and accelerate membership query in a set of items. They use modest hardware at the cost of false positives, but never produce false negatives. Signatures were traditionally used in different distributed and network applications, but in recent years their use has been extended to other fields (for instance, support for manycore/multicore parallel programming, such as data race detection, deterministic replay or transactional memory (TM)). One drawback of signatures is that they have a fixed size, and what is a good signature size for one application, may be not appropriate for another. Recently, we proposed a shared hardware module for managing signatures based on a collection of Bloom filters. It has the characteristic of hosting a variable number of signatures that change their size in runtime to adapt to the demand of the applications. However, the assignment of resources follows a single symmetric policy for all allocations leading to a module with a limited adaptability to the workloads. In this paper, we explore new techniques to allocate signatures in an asymmetric way in this module, with the aim of optimizing the resources and reducing even more the number of false positives. We explore several asymmetric strategies and their efficient hardware implementation, and we show specific examples using TM as a driver application. The results show that these strategies lead to a significant reduction in the number of false positives compared with symmetric policies. Lois Orosa 0001, Javier D. Bruguera, Elisardo Antelo |
Comput. J. | 3 |
| 2014 | Fast Radix-10 Multiplication Using Redundant BCD CodesabstractWe present the algorithm and architecture of a BCD parallel multiplier that exploits some properties of two different redundant BCD codes to speedup its computation: the redundant BCD excess-3 code (XS-3), and the overloaded BCD representation (ODDS). In addition, new techniques are developed to reduce significantly the latency and area of previous representative high-performance implementations. Partial products are generated in parallel using a signed-digit radix-10 recoding of the BCD multiplier with the digit set [-5, 5], and a set of positive multiplicand multiples (0X, 1X, 2X, 3X, 4X, 5X) coded in XS-3. This encoding has several advantages. First, it is a self-complementing code, so that a negative multiplicand multiple can be obtained by just inverting the bits of the corresponding positive one. Also, the available redundancy allows a fast and simple generation of multiplicand multiples in a carry-free way. Finally, the partial products can be recoded to the ODDS representation by just adding a constant factor into the partial product reduction tree. Since the ODDS uses a similar 4-bit binary encoding as non-redundant BCD, conventional binary VLSI circuit techniques, such as binary carry-save adders and compressor trees, can be adapted efficiently to perform decimal operations. To show the advantages of our architecture, we have synthesized a RTL model for$16\times 16$-digit and$34\times 34$-digit multiplications and performed a comparative survey of the previous most representative designs. We show that the proposed decimal multiplier has an area improvement roughly in the range 20-35 percent for similar target delays with respect to the fastest implementation. Álvaro Vázquez, Elisardo Antelo, Javier D. Bruguera |
IEEE Trans. Computers | 2 |
| 2012 | FlexSig: Implementing flexible hardware signaturesabstractWith the advent of chip multiprocessors, new techniques have been developed to make parallel programing easier and more reliable. New parallel programing paradigms and new methods of making the execution of programs more efficient and more reliable have been developed. Usually, these improvements require hardware support to avoid a system slowdown. Signatures based on Bloom filters are widely used as hardware support for parallel programing in chip multiprocessors. Signatures are used in Transactional Memory, thread-level speculation, parallel debugging, deterministic replay and other tools and applications. The main limitation of hardware signatures is the lack of flexibility: if signatures are designed with a given configuration, tailored to the requirements of a specific tool or application, it is likely that they do not fit well for other different requirements. In this paper a new hardware signature organization, called Flexible Signatures ( FlexSig ), is proposed. FlexSig can change dynamically the resources assigned to a given signature and the number of signatures in the system, by redistributing the available hardware resources according to the system requirements. This allows higher flexibility than with traditional fixed-resources signatures based on Bloom filters, while maintaining a low false positive rate. FlexSig has been evaluated by comparing it with signatures based on parallel Bloom filters, and we conclude that FlexSig outperforms (in terms of false positive rate) conventional parallel Bloom filters in most cases, due to its ability to use all the signature resources available. Lois Orosa 0001, Elisardo Antelo, Javier D. Bruguera |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Guest Editors' Introduction: Special Section on Computer ArithmeticabstractThe four articles in this special section focus on the topic of computer arithmetic and its applications. Elisardo Antelo, David Hough 0001, Paolo Ienne |
IEEE Trans. Computers | 1 |
| 2012 | Redundant Floating-Point Decimal CORDIC AlgorithmabstractIn this work, we propose a new decimal redundant CORDIC algorithm to manage transcendental functions, using floating-point representation. The algorithms determine the direction of the elementary rotation using sign estimations. Unlike binary redundant CORDIC, repetition of iterations are not required to ensure convergence since novel decimal codes have been carefully selected with sufficient redundancy to prevent any repetition. The algorithms are mapped to a low-cost unit based on a decimal 3-2 carry-save adder which can also be used as a floating-point decimal division unit. Compared to current decimal floating-point units, the implementation of our algorithm involves minor modifications of the native hardware, while providing a huge set of elementary functions. Álvaro Vázquez, Julio Villalba, Elisardo Antelo, Emilio L. Zapata |
IEEE Trans. Computers | 3 |
| 2011 | Reducing the Computation Time in (Short Bit-Width) Two's Complement MultipliersabstractTwo's complement multipliers are important for a wide range of applications. In this paper, we present a technique to reduce by one row the maximum height of the partial product array generated by a radix-4 Modified Booth Encoded multiplier, without any increase in the delay of the partial product generation stage. This reduction may allow for a faster compression of the partial product array and regular layouts. This technique is of particular interest in all multiplier designs, but especially in short bit-width two's complement multipliers for high-performance embedded cores. The proposed method is general and can be extended to higher radix encodings, as well as to any size square and m \times n rectangular multipliers. We evaluated the proposed approach by comparison with some other possible solutions; the results based on a rough theoretical analysis and on logic synthesis showed its efficiency in terms of both area and delay. Fabrizio Lamberti, Nikolaos Andrikos, Elisardo Antelo, Paolo Montuschi |
IEEE Trans. Computers | 3 |
| 2010 | Improved Design of High-Performance Parallel Decimal MultipliersabstractThe new generation of high-performance decimal floating-point units (DFUs) is demanding efficient implementations of parallel decimal multipliers. In this paper, we describe the architectures of two parallel decimal multipliers. The parallel generation of partial products is performed using signed-digit radix-10 or radix-5 recodings of the multiplier and a simplified set of multiplicand multiples. The reduction of partial products is implemented in a tree structure based on a decimal multioperand carry-save addition algorithm that uses unconventional (non BCD) decimal-coded number systems. We further detail these techniques and present the new improvements to reduce the latency of the previous designs, which include: optimized digit recoders for the generation of 2n-tuples (and 5-tuples), decimal carry-save adders (CSAs) combining different decimal-coded operands, and carry-free adders implemented by special designed bit counters. Moreover, we detail a design methodology that combines all these techniques to obtain efficient reduction trees with different area and delay trade-offs for any number of partial products generated. Evaluation results for 16-digit operands show that the proposed architectures have interesting area-delay figures compared to conventional Booth radix-4 and radix--8 parallel binary multipliers and outperform the figures of previous alternatives for decimal multiplication. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 2009 | A High-Performance Significand BCD Adder with IEEE 754-2008 Decimal RoundingabstractWe present a new method and architecture to merge efficiently IEEE 754-2008 decimal rounding with significand BCD addition and subtraction. This is a key component to improve several decimal floating-point operations such as addition, multiplication and fused multiply-add. The decimal rounding unit is based on a direct implementation of the IEEE 754-2008 rounding modes. We show that the resultant implementations for IEEE 754-2008 Decimal64 (16 precision digits) and Decimal128 (34 precision digits) formats reduce significantly the area and latency required for significand BCD addition/subtraction and decimal rounding in previous high-performance decimal floating-point adders. Álvaro Vázquez, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 2 |
| 2009 | Computation of Decimal Transcendental Functions Using the CORDIC AlgorithmabstractIn this work we propose new decimal floating-point CORDIC algorithms for transcendental function evaluation. We show how these algorithms are mapped to a state of the art Decimal Floating-Point Unit (DFPU), both considering the use of a carry--propagate adder or a carry--save redundant adder. We compared with previous decimal CORDIC proposals and with table-driven algorithms, and we concluded that our approach have significant potential advantages for transcendental function evaluation in state of the art DFPUs with minor modifications of the hardware. Álvaro Vázquez, Julio Villalba, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 3 |
| 2008 | New insights on Ling addersabstractAdders are critical for microprocessor design. Current designs use variations of parallel prefix schemes. A method introduced by Ling [7] may improve this kind of adders. However, as recent research publications demonstrate, the use of the Ling scheme in prefix adders is not a mature and clear concept. In this work we show how to easily extend any existing prefix adder topology to use the Ling method. Moreover, we use this methodology to implement the Ling scheme in a flagged prefix adder, which is an interesting building block for floating point units. Álvaro Vázquez, Elisardo Antelo |
ASAP | 2 |
| 2008 | A Low-Latency Pipelined 2D and 3D CORDIC ProcessorsabstractThe unfolded and pipelined CORDIC is a high-performance hardware element that produces a wide variety of one and two argument functions with high throughput. The reduction in delay, power, and area (cost) are of significant interest regarding this module due to its high demand for resources. The linear approximation to rotation has been proposed to achieve such reductions. However, the schemes for rotation (multiplication) and vectoring (division) complicate the implementation in a single unit. In this work, we improve the linear approximation scheme, leading to a unified implementation for rotation and vectoring, where fully parallel tree multipliers are used instead of the second half of CORDIC iterations. We also combine the linear approximation to rotation with the scale factor compensation so that the compensation is concurrently performed with the rotation process. We then extend the method to 3D CORDIC. Such an extension is not straightforward due to the lack of existing analytical expressions for the convergence of the algorithm. A comparison, using a rough area-time model and synthesis results, shows that our proposals may achieve significant reductions in delay, with no increase in area, in actual implementations. Elisardo Antelo, Julio Villalba, Emilio L. Zapata |
IEEE Trans. Computers | 1 |
| 2007 | A New Family of High.Performance Parallel Decimal MultipliersabstractThis paper introduces two novel architectures for parallel decimal multipliers. Our multipliers are based on a new algorithm for decimal carry-save multioperand addition that uses a novel BCD-4221 recoding for decimal digits. It significantly improves the area and latency of the partial product reduction tree with respect to previous proposals. We also present three schemes for fast and efficient generation of partial products in parallel. The recoding of the BCD-8421 multiplier operand into minimally redundant signed-digit radix-10, radix-4 and radix-5 representations using new recoders reduces the complexity of partial product generation. In addition, SD radix-4 and radix-5 recodings allow the reuse of a conventional parallel binary radix-4 multiplier to perform combined binary/decimal multiplications. Evaluation results show that the proposed architectures have interesting area-delay figures compared to conventional Booth radix-4 and radix-8 parallel binary multipliers and other representative alternatives for decimal multiplication. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
IEEE Symposium on Computer Arithmetic | 2 |
| 2007 | A radix-10 SRT divider based on alternative BCD codingsabstractIn this paper we present the algorithm and architecture a radix-10 floating-point divider based on an SRT non-restoring digit-by-digit algorithm. The algorithm uses conventional techniques developed to speed-up radix-2kdivision such as signed-digit (SD) redundant quotient and digit selection by constant comparison using a carry-save estimate of the partial remainder. To optimize area and latency for decimal, we include novel features such as the use of alternative BCD codings to represent decimal operands, estimates by truncation at any binary position inside a decimal digit, a single customized fast carry propagate decimal adder for partial remainder computation, initial odd multiple generation and final normalization with rounding, and register placement to exploit advanced high fanin mux-latch circuits. The rough area-delay estimations performed show that the proposed divider has a similar latency but less hardware complexity (1.3 area ratio) than a recently published high performance digit-by-digit implementation. Álvaro Vázquez, Elisardo Antelo, Paolo Montuschi |
ICCD | 2 |
| 2005 | Low Latency Digit-Recurrence Reciprocal and Square-Root Reciprocal Algorithm and ArchitectureabstractThe reciprocal and square-root reciprocal operations are important in several applications. For these operations, we present algorithms that combine a digit-by-digit module and one iteration of a quadratic-convergence approximation. The latter is implemented by a digit-recurrence, which uses the digits produced by the digit-by-digit part. In this way, both parts execute in an overlapped manner, so that the total number of cycles is about half of the number that would be required by the digit-by-digit part alone. Because of the approximation, correct rounding of the result cannot be obtained directly in all cases; we propose a variable-time implementation that produces the correctly rounded result with a small average overhead. Radix-4 implementations are described and have been synthesized. They achieve the same cycle time as the standard digit-by-digit implementation, resulting in a speed-up of about 2 and, because of the approximation part, the area factor is also about 2. We also show a combined implementation for both operations that has essentially the same complexity as that for square-root reciprocal alone. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Symposium on Computer Arithmetic | 1 |
| 2005 | Low Latency Pipelined Circular CORDICabstractThe pipelined CORDIC with linear approximation to rotation has been proposed to achieve reductions in delay, power and area; however, the schemes for rotation (multiplication) and vectoring (division) complicate implementation in a single unit. In this work, we improve the linear approximation scheme, leading to a unified implementation for rotation and vectoring where fully parallel tree multipliers are used instead of the second half of CORDIC iterations. We also combine the linear approximation to rotation with the scale factor compensation so that the compensation is performed concurrently with the rotation process. Comparison with other designs is also provided. Elisardo Antelo, Julio Villalba |
IEEE Symposium on Computer Arithmetic | 1 |
| 2005 | Digit-Recurrence Dividers with Reduced Logical DepthabstractIn this paper, we propose a class of division algorithms with the aim of reducing the delay of the selection of the quotient digit by introducing more concurrency and flexibility in its computation. From the proposed class of algorithms, we select one that moves part of the selection function out of the critical path, with a corresponding reduction in the critical path compared with existing alternatives: we present the algorithm and describe the architectures for radix 4 and for radix 16. For radix 16, we use the scheme of overlapping two radix-4 stages. In both cases, radix 4 and radix 16, we show that our algorithms allow the design of units with well-balanced critical paths with consequent decreases of the cycle times. Moreover, in the radix-16 case, we include some additional speculation techniques. To estimate the speedup, we used a rough timing model based on logical effort. For both radices, we estimate a speedup of about 25 percent with respect to previous implementations. In the radix-4 case, this is achieved by using roughly the same area, while, in the radix-16 case, the area is increased by about 30 percent. We verified our estimations by performing a synthesis of the radix-4 units. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
IEEE Trans. Computers | 1 |
| 2005 | High-Throughput CORDIC-Based Geometry Operations for 3D Computer GraphicsabstractGraphics processors require strong arithmetic support to perform computational kernels over data streams. Because of the current implementation using the basic arithmetic operations, the algorithms are given in algebraic terms. However, since the operations are really of a geometric nature, it seems to us that more flexibility in the implementation is obtained if the description is given in a high-level geometrical form. As a consequence of this line of thought, this paper is an attempt to reconsider some kernels in a graphics processor to obtain implementations that are potentially more scalable than just replicating the modules used in conventional implementations. We present the formulation of representative 3D computer graphics operations in terms of CORDIC-type primitives. Then, we briefly outline a stream processor based on CORDIC-type modules to efficiently implement these graphic operations. We perform a rough comparison with current implementations and conclude that the CORDIC-based alternative might be attractive. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 2 |
| 2003 | Radix-4 Reciprocal Square-Root and Its Combination with Division and Square RootabstractIn this work, we present a reciprocal square root algorithm by digit recurrence and selection by a staircase function and the radix-4 implementation. As in similar algorithms for division and square root, the results are obtained correctly rounded in a straightforward manner (in contrast to existing methods to compute the reciprocal square root). Although, apparently, a single selection function can only be used for j /spl ges/ 2 (the selection constants are different for j = 0, j = 1, and j /spl ges/ 2), we show that it is possible to use a single selection function for all iterations. We perform a rough comparison with existing methods and we conclude that our implementation is a low hardware complexity solution with moderate latency, especially for exactly rounded results. We also extend the unit to support division and square root with the same selection function and with slight modifications in the initialization of the reciprocal square root unit. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 2 |
| 2002 | Fast Radix-4 Retimed Division with Selection by ComparisonsabstractSince a large portion of the critical path in an implementation of radix-4 division corresponds to the delay of the quotient-digit selection module, it is of interest to reduce this delay. The proposal of this paper extends the approach presented recently of prestoring the selection constants corresponding to the actual value of the divisor and to perform the determination of the quotient digit by carry-free subtraction and sign detection. This extension consists in advancing the subtraction so that it is outside of the critical path. This advancement also provides the possibility of placing the registers so as to minimize the cycle time. We present the method and report results of synthesis using a family of standard cells. We conclude that the extension results in a speedup of 1.35 with respect to the basic implementation and of 1.3 with respect to the previously mentioned approach. We estimate that the areas of all three units are about the same. Elisardo Antelo, Tomás Lang, Paolo Montuschi, Alberto Nannarelli |
ASAP | 1 |
| 2001 | Correctly Rounded Reciprocal Square-Root by Digit Recurrence and Radix-4 ImplementationabstractWe present a reciprocal square-root algorithm by digit recurrence and selection by a staircase function, and the radix-4 implementation. As similar algorithms for division and square-root, the results are obtained correctly rounded in a straightforward manner (in contrast to existing methods to compute the reciprocal square-root). Although apparently a single selection function can only be used for j/spl ges/2 (the selection constants are different for j=0, j=1 and j/spl ges/2), we show that it is possible to use a single selection function for all iterations. We perform a rough comparison with existing methods and we conclude that our implementation is a low hardware complexity solution with moderate latency, specially for exactly rounded results. Tomás Lang, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 2 |
| 2000 | Very-High Radix Circular CORDIC: Vectoring and Unified Rotation/VectoringabstractA very-high radix algorithm and implementation for circular CORDIC is presented. We first present in depth the algorithm for the vectoring mode in which the selection of the digits is performed by rounding of the control variable. To assure convergence with this kind of selection, the operands are prescaled. However, in the CORDIC algorithm, the coordinate x varies during the execution so several scalings might be needed; we show that two scalings are sufficient. Moreover, the compensation of the variable scale factor (including the CORDIC scale factor and the prescaling factors) is done by computing the logarithm of the scale factor and performing the compensation by an exponential. Then, we combine, in a unified unit, the proposed vectoring algorithm and the very-high radix rotation algorithm, which was previously proposed by the authors. We compare with low-radix implementations in terms of latency and hardware complexity. Estimations of the delay for 32-bit precision show a speedup of about two with respect to the radix-4 case with redundant addition. This speedup is obtained at the cost of an increase in the hardware complexity, which is moderate for the pipelined implementation. We also compare at the algorithmic level with other very-high radix proposals, demonstrating the advantages of our algorithms. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Trans. Computers | 1 |
| 1999 | Very-High Radix CORDIC Vectoring with Scalings and Selection by RoundingabstractA very-high radix algorithm and implementation for circular CORDIC in vectoring mode is presented. As for division, to simplify the selection function, the operands are pre-scaled. However in the CORDIC algorithm the coordinate x varies during the execution so several scalings might be needed; we show that two scalings are sufficient. Moreover, the compensation of the variable scale factor is done by computing the logarithm of the scale factor and performing the compensation by an exponential. Estimations of the delay for 32 bit precision show a speed up of about two with respect to the radix-4 case with redundant addition. This speed up is obtained at the cost of an increase in the hardware complexity, which is moderate for the pipelined implementation. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Symposium on Computer Arithmetic | 1 |
| 1998 | Computation of sqrt(x/d) in a Very High Radix Combined Division/Square-Root Unit with ScalingabstractA very-high radix digit-recurrence algorithm for the operation /spl radic/(x/d) is developed, with residual scaling and digit selection by rounding. This is an extension of the division and square-root algorithms presented previously, and for which a combined unit was shown to provide a fast execution of these operations. The architecture of a combined unit to execute division, square-root, and /spl radic/(x/d) is described, with inverse square-root as a special case. A comparison with the corresponding combined division and square-root unit shows a similar cycle time and an increase of one cycle for the extended operation with respect to square-root. To obtain an exactly rounded result for the extended operation a datapath of about 2n bits is needed. An alternative is proposed which requires approximately the same width as for square-root, but produces a result with an error of less than one ulp. The area increase with respect to the division and square root unit should be no greater than 15 percent. Consequently, whenever a very high radix unit for division and square-root seems suitable, it might be profitable to implement the extended unit instead. Elisardo Antelo, Tomás Lang, Javier D. Bruguera |
IEEE Trans. Computers | 1 |
| 1998 | CORDIC Vectoring with Arbitrary Target ValueabstractThe computation of additional functions in the CORDIC module increases its flexibility. We consider here the extension of the vectoring mode (angle calculation) so that the vector is rotated until one of the coordinates (for instance y) attains a target value t (in contrast to the value 0, as in standard vectoring). The main problem in the algorithm is that the modulus of the vector is scaled in each CORDIC iteration so that a direct comparison of y[j] with t does not assure convergence. We present a scheme that overcomes this and in which the implementation consists of a standard CORDIC module plus a module to determine the direction of rotation. This improves over a previous proposal in which more complex iterations are introduced as part of the CORDIC algorithm. Moreover, an error analysis is performed to determine the datapath width required for convergence. Since this width is large, we consider also the characteristics of the algorithm for a narrower datapath. Tomás Lang, Elisardo Antelo |
IEEE Trans. Computers | 2 |
| 1998 | A novel design of a two operand normalization circuitabstractThis paper presents a new design for two operand normalization. The two operand normalization operation involves the normalization of at least one of two operands by left shifting both by the same amount. Our design performs the computation of the shift by making an OR of the bits of both operands in a tree network, encoding the position of the first nonzero bit. The encoded position is obtained most significant bit first, and then there is an overlapping with the shifting operation. The design we propose replaces two leading zero detector circuits and a comparator, that are present in the conventional approach. Our scheme demonstrates to be more area efficient than the conventional one. The circuit we propose is useful in floating point complex multiplication and COordinate Rotation DIgital Computer (CORDIC) processors. Elisardo Antelo, Montserrat Bóo, Javier D. Bruguera, Emilio L. Zapata |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1997 | CORDIC Vectoring with Arbitrary Target ValueabstractThe computation of additional functions in the CORDIC module increases its flexibility. We consider the extension of the vectoring mode (angle calculation) so that the vector is rotated until one of the coordinates (for instance, y) attains a target value t (in contrast to the value 0, as in standard vectoring). The main problem in the algorithm is that the modulus of the vector is scaled in each CORDIC iteration so that a direct comparison of y[j] with t does not assure convergence. We present a scheme that overcomes this and in which the implementation consists of a standard CORDIC module plus a module to determine the direction of rotation. This improves over a previous proposal in which more complex iterations are introduced as part of the CORDIC algorithm. Tomás Lang, Elisardo Antelo |
IEEE Symposium on Computer Arithmetic | 2 |
| 1997 | CORDIC-based computation of arccos and arcsinabstractCORDIC-based algorithms to compute cos/sup -1/(t), sin/sup -1/(t) and /spl radic/(1-t/sup 2/) are proposed. The implementation requires a standard CORDIC module plus a module to compute the direction of rotation, this being the same hardware required for the extended CORDIC vectoring, recently proposed by the authors. Although these functions can be obtained as a special case of this extended vectoring, the specific algorithm we propose here presents two significant improvements: (1) it achieves an angle granularity of 2/sup -n/ using the same datapath width as the standard CORDIC algorithm (about n bits, instead of about 2n which would be required using the extended veetoring), and (2) no repetitions of iterations are needed. The proposed algorithm is compatible with the extended vectoring and, in contrast with previous implementations, the number of iterations and the delay of each iteration are the same as for the conventional CORDIC algorithm. Tomás Lang, Elisardo Antelo |
ASAP | 2 |
| 1997 | Error Analysis and Reduction for Angle Calculation Using the CORDIC AlgorithmabstractIn this paper, we consider the errors appearing in angle computations with the CORDIC algorithm (circular and hyperbolic coordinate systems) using fixed-point arithmetic. We include errors arising not only from the finite number of iterations and the finite width of the data path, but also from the finite number of bits of the input. We show that this last contribution is significant when both operands are small and that the error is acceptable only if an input normalization stage is included, making unsatisfactory other previous proposals to reduce the error. We propose a method based on the prescaling of the input operands and a modified CORDIC recurrence and show that it is a suitable alternative to the input normalization with a smaller hardware cost. This solution can also be used in pipelined architectures with redundant carry-save arithmetic. Elisardo Antelo, Javier D. Bruguera, Tomás Lang, Emilio L. Zapata |
IEEE Trans. Computers | 1 |
| 1997 | High Performance Rotation Architectures Based on the Radix-4 CORDIC AlgorithmabstractTraditionally, CORDIC algorithms have employed radix-2 in the first n/2 microrotations (n is the precision in bits) in order to preserve a constant scale factor. The authors present a full radix-4 CORDIC algorithm in rotation mode and circular coordinates and its corresponding selection function, and propose an efficient technique for the compensation of the nonconstant scale factor. Three radix-4 CORDIC architectures are implemented: 1) a word serial architecture based on the zero skipping technique, 2) a pipelined architecture, and 3) an application specific architecture (the angles are known beforehand). The first two are general purpose implementations where redundant (carry-save) or nonredundant arithmetic can be used, whereas the last one is a simplification of the first two. The proposed architectures present a good trade-off between latency and hardware complexity when compared with existing CORDIC architectures. Elisardo Antelo, Julio Villalba, Javier D. Bruguera, Emilio L. Zapata |
IEEE Trans. Computers | 1 |
| 1996 | Radix-4 Vectoring Cordic Algorithm And ArchitecturesabstractIn this paper we present a new CORDIC algorithm for the vectoring mode, based on the use of radix-4 preserving a complexity in the microrotations that is similar to that of the conventional radix-2 CORDIC. The use of this radix, together with the inclusion in the CORDIC algorithm of the zero skipping technique, reduces by more than half the number of iterations with respect to the conventional radix 2 CORDIC, with the consequent reduction of time in recursive architectures or area in pipelined architectures. In processes such as SVD or matrix triangularization in which the evaluation of the rotation angle is required, this algorithm is shown to be specially efficient. Julio Villalba, J. C. Arrabal, Emilio L. Zapata, Elisardo Antelo, Javier D. Bruguera |
ASAP | 4 |
| 1996 | Unified Mixed Radix 2-4 Redundant CORDIC ProcessorabstractWe present a unified mixed radix CORDIC algorithm with carry-save arithmetic with a constant scale factor. The pipelined architecture of the processor is determined by a unique sequence of microrotations for the two modes of operation (rotation and vectoring) in circular and hyperbolic coordinates. The combination of radix-2 and radix-4 microrotations allows us to reduce the latency and size of the pipeline significantly. The unified algorithm is based on the correcting microrotation method, which we have extended to the vectoring mode in hyperbolic coordinates. We have also generalized the use of radix-4 microrotations to the two operation modes and coordinate systems. Elisardo Antelo, Javier D. Bruguera, Emilio L. Zapata |
IEEE Trans. Computers | 1 |
| 1995 | Redundant CORDIC Rotator Based on Parallel PredictionabstractWe present a Cordic rotator, using carry-save arithmetic, based on the prediction of all the coefficients into which the rotation angle is decomposed. The prediction algorithm is based on the use of radix-2 microrotations with multiple shifts in the first iterations and the use of a redundant radix-2 and radix-4 representation for the coefficients in the rest of the microrotations. The use of multiple shifts facilitates the prediction of the coefficients in the case of microrotations where i/spl les/n/4, being n the precision of the algorithm, and the use of radix-4 microrotations helps to reduce the total number of iterations. The prediction is carried out using the redundant representation of the z coordinate, without any need for conversions to a non-redundant representation. Finally, we present a VLSI architecture based on this algorithm. As the production of the coefficients is very fast, and they are known before starting each microrotation, the resulting architecture can be highly pipelined and consequently appropriate for applications where high speeds are required.> Elisardo Antelo, Javier D. Bruguera, Julio Villalba, Emilio L. Zapata |
IEEE Symposium on Computer Arithmetic | 1 |
| 1995 | Digit On-line Large Radix CORDIC RotatorabstractMany applications figure the evaluation of rotations at high speeds. However there is a trade-off between the chip area and the latency. In this paper we develop a digit on-line pipelined array architecture based on the radix-4 CORDIC algorithm in rotation mode. The radix-4 CORDIC algorithm halves the number of microrotations with respect the traditionally radix-2 algorithm with the drawback of a non-constant scale factor. Seeking a good compromise between silicon area and latency we have used digit on-line processing. This way the data inputs the processor in blocks of bits (digits) in MSD-first mode of processing. We have used redundant carry-save arithmetic to allow carry-free additions and on-line processing. The designed processor demonstrates to have a better performance than previous digit on-line architectures. Roberto R. Osorio, Elisardo Antelo, Javier D. Bruguera, Julio Villalba, Emilio L. Zapata |
ASAP | 2 |
| 1995 | CORDIC Architectures with Parallel Compensation of the Scale FactorabstractThe compensation of scale factor imposes significant computation overhead on the CORDIC algorithm. In this paper we will propose two algorithms and architectures in order to perform the compensation of the scale factor in parallel with the computation of the CORDIC iterations. This way it is not necessary to carry out the final multiplication or add scaling iterations in order to achieve the compensation. With the architectures we propose the dependence on n of the compensation of the scale factor disappears, and this considerably reduces the latency of the system. The architectures developed are optimized solutions for the different operating modes of the CORDIC both in conventional and in redundant arithmetic. Julio Villalba, José A. Hidalgo-López, Emilio L. Zapata, Elisardo Antelo, Javier D. Bruguera |
ASAP | 4 |
| 1993 | Design of a Pipelined Radix 4 CORDIC Processor
Javier D. Bruguera, Elisardo Antelo, Emilio L. Zapata |
Parallel Comput. | 2 |