Julio Villalba

dblp:31/4679 · also Julio Villalba-Moreno · DBLP profile ↗
← Back
38ranked-venue papers
13as first author
0since 2021 · last 2020
0000-0001-8557-3876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 11 first-authorTheory of computation · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Integrated circuit design · 60% Processor architecture and microarchitecture · 36% Reconfigurable computing and FPGAs · 4%
Computer graphics and multimedia
1 paper
Image and video processing · 67% Multimedia analysis and retrieval · 33%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.652016
New Formats for Computing with Real-Numbers under Round-to-Nearest · IEEE Trans. Computers 2016
Decimal Multiformat Online Addition · IEEE Trans. Computers 2016
Multioperand Redundant Adders on FPGAs · IEEE Trans. Computers 2013
Processor architecture and microarchitecture
computer arithmetic
0.552016
New Formats for Computing with Real-Numbers under Round-to-Nearest · IEEE Trans. Computers 2016
Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012
A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008
Integrated circuit design › digital circuit design
arithmetic circuit design
0.422016
Decimal Multiformat Online Addition · IEEE Trans. Computers 2016
Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012
Processor architecture and microarchitecture › computer arithmetic
floating-point arithmetic
0.422018
Unbiased Rounding for HUB Floating-Point Addition · IEEE Trans. Computers 2018
Double-Residue Modular Range Reduction for Floating-Point Hardware Implementations · IEEE Trans. Computers 2006
Integrated circuit design › digital arithmetic circuits › floating-point unit design
floating-point adder
0.312018
Unbiased Rounding for HUB Floating-Point Addition · IEEE Trans. Computers 2018
Integrated circuit design › digital arithmetic circuits › floating-point unit design
rounding
0.312018
Unbiased Rounding for HUB Floating-Point Addition · IEEE Trans. Computers 2018
Integrated circuit design › digital arithmetic circuits › decimal arithmetic
decimal adder
0.212016
Decimal Multiformat Online Addition · IEEE Trans. Computers 2016
Processor architecture and microarchitecture › computer arithmetic
floating-point representation
0.212016
New Formats for Computing with Real-Numbers under Round-to-Nearest · IEEE Trans. Computers 2016
Integrated circuit design › digital arithmetic circuits
CORDIC
0.232012
Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012
A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008
High Performance Rotation Architectures Based on the Radix-4 CORDIC Algorithm · IEEE Trans. Computers 1997
Processor architecture and microarchitecture › computer arithmetic
multioperand addition
0.222013
Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012
Multioperand Redundant Adders on FPGAs · IEEE Trans. Computers 2013
Reconfigurable computing and FPGAs
FPGA arithmetic
0.212013
Multioperand Redundant Adders on FPGAs · IEEE Trans. Computers 2013
Integrated circuit design › digital arithmetic circuits
redundant adders
0.212013
Multioperand Redundant Adders on FPGAs · IEEE Trans. Computers 2013
Integrated circuit design › digital arithmetic circuits
online arithmetic
0.112012
Radix-2 Multioperand and Multiformat Streaming Online Addition · IEEE Trans. Computers 2012
Integrated circuit design › digital arithmetic circuits
special function unit
0.112012
Redundant Floating-Point Decimal CORDIC Algorithm · IEEE Trans. Computers 2012
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add
0.112018
Unbiased Rounding for HUB Floating-Point Addition · IEEE Trans. Computers 2018
Processor architecture and microarchitecture › computer arithmetic
elementary function evaluation
0.112006
Double-Residue Modular Range Reduction for Floating-Point Hardware Implementations · IEEE Trans. Computers 2006
Processor architecture and microarchitecture › arithmetic unit
arithmetic unit design
0.022008
A Low-Latency Pipelined 2D and 3D CORDIC Processors · IEEE Trans. Computers 2008
High Performance Rotation Architectures Based on the Radix-4 CORDIC Algorithm · IEEE Trans. Computers 1997
Image and video processing › feature detection
hough transform
0.011995
A fast Hough transform for segment detection · IEEE Trans. Image Process. 1995
Multimedia analysis and retrieval
image analysis
0.011995
A fast Hough transform for segment detection · IEEE Trans. Image Process. 1995
Image and video processing › feature extraction
line segment extraction
0.011995
A fast Hough transform for segment detection · IEEE Trans. Image Process. 1995

Methods — techniques the papers use, named apart from their topics

error analysis · 0.3synthesis · 0.2radix-β number systems · 0.2code conversion · 0.2carry propagation avoidance · 0.2carry-chain resource utilization · 0.2HDL parameterization · 0.2sign estimation · 0.1carry-save representation · 0.1carry-save addition · 0.1parallelization · 0.0
YearPublicationVenuePosition
2020 Floating-Point Fused Multiply-Add under HUB Format
abstract
The Half-Unit-Biased (HUB) format has interesting advantages for implementing floating-point arithmetic which has been proved for the four basic arithmetic operations as well as square root. Nevertheless, although Floating-point Fused Multiply-add (FMA) operation (AxB + C) is one of the most important and complex arithmetic instructions in modern processors, FMA operation for HUB numbers has not been confronted yet. In this paper, we present a design to deal with this operation under HUB format. The key points to turn the conventional FMA architecture into a HUB unit are explained. Comparing the ASIC implementation of a HUB FMA unit with the conventional one, the former reduces the required area and power up to 38% and 35%, respectively, for single-precision. For BFloat16, the HUB FMA increases the speed a 15%, and even then, reduces the area and power by 26% and 12%, respectively.
Javier Hormigo, Julio Villalba, Sonia Gonzalez-Navarro
ARITH2
2019 Reproducible Summation Under HUB Format
abstract
Floating point reproducibility is a property claimed by programmers and end users. Half-Unit-Biased (HUB) is a new representation format in which the round to nearest is carried out by truncation, preventing any carry propagation and saving time and area. In this paper we study the reproducible summation of HUB numbers by using a error-free vector transformation technique, providing both a specific architecture and the usage of combined HUB/Standard floating point adders to achieve a reproducible result.
Julio Villalba, Javier Hormigo, Francisco J. Jaime
ARITH1
2018 Unbiased Rounding for HUB Floating-Point Addition
abstract
Half-Unit-Biased (HUB) is an emerging format based on shifting the represented numbers by half Unit in the Last Place. This format simplifies two's complement and round-to-nearest operations by preventing any carry propagation. This saves power consumption, time and area. Taking into account that the IEEE floating-point standard uses an unbiased rounding as the default mode, this feature is also desirable for HUB approaches. In this paper, we study the unbiased rounding for HUB floating-point addition in both as standalone operation and within FMA. We show two different alternatives to eliminate the bias when rounding the sum results, either partially or totally. We also present an error analysis and the implementation results of the proposed architectures to help the designers to decide what their best option are.
Julio Villalba, Javier Hormigo, Sonia Gonzalez-Navarro
IEEE Trans. Computers1
2017 Floating Point Square Root under HUB Format
abstract
Unit-Biased (HUB) is an emerging format based on shifting the representation line of the binary numbers by half unit in the last place. The HUB format is specially relevant for computers where rounding to nearest is required because it is performed simply by truncation. From a hardware point of view, the circuits implementing this representation save both area and time since rounding does not involve any carry propagation. Designs to perform the four basic operations have been proposed under HUB format recently. Nevertheless, the square root operation has not been confronted yet. In this paper we present an architecture to carry out the square root operation under HUB format for floating point numbers. The results of this work keep supporting the fact that the HUB representation involves simpler hardware than its conventional counterpart for computers requiring round-to-nearest mode.
Julio Villalba, Javier Hormigo
ICCD1
2017 Introduction to the Special Issue on Computer Arithmetic
abstract
The papers in this special issue focus on computer arithmetic which is used in many applications, usually totally silently (one should keep in mind that even when running programs that are not at all numeric, memory addresses are computed, which involves additions, multiplications, and sometimes divisions). However, in some areas, it plays a central role.
Javier Hormigo, Jean-Michel Muller, Stuart F. Oberman, Nathalie Revol, Arnaud Tisserand, Julio Villalba
IEEE Trans. Computers6
2016 Digit Recurrence Floating-Point Division under HUB Format
abstract
Half-Unit-Biased format is based on shifting the representation line of the binary numbers by half Unit in theLast Place. The main feature of this format is that the roundto nearest is carried out by a simple truncation, preventing any carry propagation and saving time and area. Algorithms and architectures have been defined for ddition/substraction and multiplication operations under this format. Nevertheless, the division operation has not been confronted yet. In this paper we deal with the floating-point division under HUB format, studying the architecture for the digit recurrence method, including the on-the-fly conversion of the signed digit quotient.
Julio Villalba
ARITH1
2016 Decimal Multiformat Online Addition
abstract
This paper presents and analyzes two different strategies for designing multiformat online decimal adders ($\text{olDFA}_{\text{Mformat}}$). The first strategy uses a code conversion stage plus an online Decimal Full Adder (olDFA); the second one involves designing specific adders by modifying the architecture of the olDFA. These strategies are applied in the design of specific architectures to deal with financial analysis calculations. We use synthesis results to verify the theoretical aspects of the designs and to analyze the robustness and lacks of the strategies. The guidelines presented in the paper are valuable to designers of online multiformat-based solutions.
Carlos Garcia-Vega, Sonia Gonzalez-Navarro, Pedro Balboa-La Chica, Julio Villalba
IEEE Trans. Computers4
2016 New Formats for Computing with Real-Numbers under Round-to-Nearest
abstract
In this paper, a new family of formats to deal with real number for applications requiring round to nearest is proposed. They are based on shifting the set of exactly represented numbers which are used in conventional radix-$\beta$number systems. This technique allows performing radix complement and round to nearest without carry propagation with negligible time and hardware cost. Furthermore, the proposed formats have the same storage cost and precision as standard ones. Since conversion to conventional formats simply require appending one extra-digit to the operands, standard circuits may be used to perform arithmetic operations with operands under the new format. We also extend the features of the RN-representation system and carry out a thorough comparison between both representation systems. We conclude that the proposed representation system is generally more adequate to implement systems for computation with real number under round-to-nearest.
Javier Hormigo, Julio Villalba
IEEE Trans. Computers2
2016 Measuring Improvement When Using HUB Formats to Implement Floating-Point Systems Under Round-to-Nearest
abstract
This paper analyzes the benefits of using half-unit-biased (HUB) formats to implement floating-point (FP) arithmetic under a round-to-nearest mode from a quantitative point of view. Using the HUB formats to represent numbers allows the removal of the rounding logic of arithmetic units, including sticky-bit computation. This is shown for FP adders, multipliers, and converters. Experimental analysis demonstrates that the HUB formats and the corresponding arithmetic units maintain the same accuracy as the conventional ones. On the other hand, the implementation of these units, based on basic architectures, shows that the HUB formats simultaneously improve area, speed, and power consumption. In addition, based on the data obtained from the synthesis, an HUB single-precision adder is ~14% faster but consumes 38% less area and 26% less power than the conventional adder. Similarly, an HUB single-precision multiplier is 17% faster, uses 22% less area, and consumes slightly less power than the conventional multiplier. At the same speed, the adder and the multiplier achieve area and power reductions of up to 50% and 40%, respectively.
Javier Hormigo, Julio Villalba
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Efficient floating-point representation for balanced codes for FPGA devices
abstract
We propose a floating-point representation to deal efficiently with arithmetic operations in codes with a balanced number of additions and multiplications for FPGA devices. The variable shift operation is very slow in these devices. We propose a format that reduces the variable shifter penalty. It is based on a radix-64 representation such that the number of the possible shifts is considerably reduced. Thus, the execution time of the floating-point addition is highly optimized when it is performed in an FPGA device, which compensates for the multiplication penalty when a high radix is used, as experimental results have shown. Consequently, the main problem of previous specific high-radix FPGA designs (no speedup for codes with a balanced number of multiplications and additions) is overcome with our proposal. The inherent architecture supporting the new format works with greater bit precision than the corresponding single precision (SP) IEEE-754 standard.
Julio Villalba, Javier Hormigo, Francisco Corbera, Mario A. González, Emilio L. Zapata
ICCD1
2013 Multioperand Redundant Adders on FPGAs
abstract
Although redundant addition is widely used to design parallel multioperand adders for ASIC implementations, the use of redundant adders on Field Programmable Gate Arrays (FPGAs) has generally been avoided. The main reasons are the efficient implementation of carry propagate adders (CPAs) on these devices (due to their specialized carry-chain resources) as well as the area overhead of the redundant adders when they are implemented on FPGAs. This paper presents different approaches to the efficient implementation of generic carry-save compressor trees on FPGAs. They present a fast critical path, independent of bit width, with practically no area overhead compared to CPA trees. Along with the classic carry-save compressor tree, we present a novel linear array structure, which efficiently uses the fast carry-chain resources. This approach is defined in a parameterizable HDL code based on CPAs, which makes it compatible with any FPGA family or vendor. A detailed study is provided for a wide range of bit widths and large number of operands. Compared to binary and ternary CPA trees, speedups of up to 2.29 and 2.14 are achieved for 16-bit width and up to 3.81 and 3.11 for 64-bit width.
Javier Hormigo, Julio Villalba, Emilio L. Zapata
IEEE Trans. Computers2
2012 On-line Decimal Adder with RBCD Representation
abstract
In this paper we present the design of an on-line adder dealing with two RBCD numbers. This basic element is intended to be be used in any on-line system in which the addition is involved. We obtain the on-line adder by serialization of a recent parallel RBCD adder with minimum latency.To reduce the cycle time a pipelined version is proposed. To deal with data stream the throughput has been reduced to its theoretical minimum possible value by a negligible cost hardware modification. Finally, actual implementation results for 16 digits (i.e. decimal64 format) are presented.
Carlos Garcia-Vega, Sonia Gonzalez-Navarro, Julio Villalba, Emilio L. Zapata
ASAP3
2012 Redundant Floating-Point Decimal CORDIC Algorithm
abstract
In this work, we propose a new decimal redundant CORDIC algorithm to manage transcendental functions, using floating-point representation. The algorithms determine the direction of the elementary rotation using sign estimations. Unlike binary redundant CORDIC, repetition of iterations are not required to ensure convergence since novel decimal codes have been carefully selected with sufficient redundancy to prevent any repetition. The algorithms are mapped to a low-cost unit based on a decimal 3-2 carry-save adder which can also be used as a floating-point decimal division unit. Compared to current decimal floating-point units, the implementation of our algorithm involves minor modifications of the native hardware, while providing a huge set of elementary functions.
Álvaro Vázquez, Julio Villalba, Elisardo Antelo, Emilio L. Zapata
IEEE Trans. Computers2
2012 Radix-2 Multioperand and Multiformat Streaming Online Addition
abstract
In this paper, we present multioperand radix-2 online addition using different data representations (signed-digit, two's complement, and carry-save), in particular cases in which operands with different representations are added. We use the previously defined online full adder (olFA) as a component to build different multioperand online architectures. To merge data with different representations, an inner conversion of data is performed, eliminating any conversion stage and penalty time. We propose a technique to build multioperand trees efficiently and give six practical rules to deal with different kinds of data in the same adder. For addition of a stream of data, we determine the minimum number of separation cycles required to isolate two successive computations and propose a novel hardware technique that eliminates completely the separation cycles, resulting in the maximum throughput possible.
Julio Villalba, Tomás Lang, Javier Hormigo
IEEE Trans. Computers1
2011 High-Speed Algorithms and Architectures for Range Reduction Computation
abstract
Range reduction is a crucial step for accuracy in trigonometric functions evaluation. This paper shows and compares a set of algorithms for additive range reduction computation and their corresponding application-specific integrated circuit implementations (ensuring an accuracy of one unit in the last place). A word-serial architecture implementation has been used as a reference for clearer comparisons. Besides, a new table-based pipelined architecture for range reduction has also been proposed.
Francisco J. Jaime, M. A. Sánchez, Javier Hormigo, Julio Villalba, Emilio L. Zapata
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Computation of Decimal Transcendental Functions Using the CORDIC Algorithm
abstract
In this work we propose new decimal floating-point CORDIC algorithms for transcendental function evaluation. We show how these algorithms are mapped to a state of the art Decimal Floating-Point Unit (DFPU), both considering the use of a carry--propagate adder or a carry--save redundant adder. We compared with previous decimal CORDIC proposals and with table-driven algorithms, and we concluded that our approach have significant potential advantages for transcendental function evaluation in state of the art DFPUs with minor modifications of the hardware.
Álvaro Vázquez, Julio Villalba, Elisardo Antelo
IEEE Symposium on Computer Arithmetic2
2009 Efficient Implementation of Carry-Save Adders in FPGAs
abstract
Most field programmable gate array (FPGA) devices have a special fast carry propagation logic intended to optimize addition operations. The redundant adders do not easily fit into this specialized carry-logic and, consequently, they require double hardware resources than carry propagate adders, while showing a similar delay for small size operands. Therefore, carry-save adders are not usually implemented on FPGA devices, although they are very useful in ASIC implementations. In this paper we study efficient implementations of carry-save adders on FPGA devices, taking advantage of the specialized carry-logic. We show that it is possible to implement redundant adders with a hardware cost close to that of a carry propagate adder. Specifically, for 16 bits and bigger wordlengths, redundant adders are clearly faster and have an area requirement similar to carry propagate adders. Among all the redundant adders studied, the 4:2 compressor is the fastest one, presents the best exploitation of the logic resources within FPGA slices and the easiest way to adapt classical algorithms to efficiently fit FPGA resources.
Javier Hormigo, Manuel Ortiz, Francisco Javier Quiles-Latorre, Francisco J. Jaime, Julio Villalba, Emilio L. Zapata
ASAP5
2008 SIMD Enhancements for a Hough Transform Implementation
abstract
The Hough transform is a line detection algorithm widely used within image processing applications, showing several variations depending on the shape which is intended to be detected. This paper provides for some new SIMD instructions aided by a specialized look-up table specifically devised for an improved Hough transform algorithm implementation, although they may also be used within other algorithms implementation with a slight modification. The new approach has been tested and evaluated using SimpleScalar tool set.
Francisco J. Jaime, Javier Hormigo, Julio Villalba, Emilio L. Zapata
DSD3
2008 New SIMD instructions set for image processing applications enhancement
abstract
Due to its inherent data parallelism, image processing applications benefit from multimedia extensions SIMD instructions within general purpose processors. However, current multimedia extensions do not allow to simultaneously address different memory positions using indirect addressing, i. e. the desired memory position address is located within a register. This restriction forces to sequential execution at some program points. This paper shows a new set of instructions providing parallel indirect addressing to a specialized table and intended to be added to existing multimedia extensions. In order to evaluate the new instructions usefulness and feasibility, we have used SimpleScalar for testing some image processing applications, getting in some cases a speed up of 3.5.
Francisco J. Jaime, Javier Hormigo, Julio Villalba, Emilio L. Zapata
ICIP3
2008 A Low-Latency Pipelined 2D and 3D CORDIC Processors
abstract
The unfolded and pipelined CORDIC is a high-performance hardware element that produces a wide variety of one and two argument functions with high throughput. The reduction in delay, power, and area (cost) are of significant interest regarding this module due to its high demand for resources. The linear approximation to rotation has been proposed to achieve such reductions. However, the schemes for rotation (multiplication) and vectoring (division) complicate the implementation in a single unit. In this work, we improve the linear approximation scheme, leading to a unified implementation for rotation and vectoring, where fully parallel tree multipliers are used instead of the second half of CORDIC iterations. We also combine the linear approximation to rotation with the scale factor compensation so that the compensation is concurrently performed with the rotation process. We then extend the method to 3D CORDIC. Such an extension is not straightforward due to the lack of existing analytical expressions for the convergence of the algorithm. A comparison, using a rough area-time model and synthesis results, shows that our proposals may achieve significant reductions in delay, with no increase in area, in actual implementations.
Elisardo Antelo, Julio Villalba, Emilio L. Zapata
IEEE Trans. Computers2
2007 Improving the Throughput of On-line Addition for Data Streams
abstract
In this paper we deal with the throughput of on–line addition for a stream of data. This throughput is directly related to the initiation interval between two successive instances. The on–line delay for the addition of two signed–digit (or carry–save) numbers is two, and N+2 cycles are classically used to compute a new pair of N–digit data (initiation interval: N+2). In this paper we present some techniques to reduce the initiation interval to N (which is the theoretical minimum value) with a very small amount of hardware or N+1 with no hardware cost. For short operands, this might have a significant effect on the throughput.
Julio Villalba, Javier Hormigo, Tomás Lang
ASAP1
2006 Pipelined Range Reduction for Floating Point Numbers
abstract
This paper presents a new pipelined architecture to deal with range reduction for floating point representation. It is based on Horner's scheme and a look-up table. The overall design has been optimized for a module equal to 2π, which is the most widely used due to trigonometric functions requirements. To ensure an accuracy of one unit in the last place (ULP), a complete error propagation study has been carried out.
Francisco J. Jaime, Julio Villalba, Javier Hormigo, Emilio L. Zapata
ASAP2
2006 Fast Full-Search Block Matching Algorithm Motion Estimation Alternatives in FPGA
abstract
Block matching motion estimation takes a great part of the processing time for video encoding. To accelerate this process is must to reach real time video coding. The best motion vector is obtained by full-search block matching algorithm which has to be usually implemented by hardware. In recent years, several FPGA based designs have been proposed since these devices support high number of process elements in parallel mode. In this paper a survey of recent architectures to perform the full-search block matching algorithm in FPGAs is presented. A further comparison on terms of frames per second reached, hardware cost in CLB slices and system frequency is presented
Joaquín Olivares 0001, José Ignacio Benavides Benítez, Javier Hormigo, Julio Villalba, Emilio L. Zapata
FPL4
2006 Double-Residue Modular Range Reduction for Floating-Point Hardware Implementations
abstract
In this paper, we present a novel algorithm and the corresponding architecture for performing range reduction, which is a preprocessing task required for the evaluation of some elementary functions such as trigonometric and exponential-based functions. The proposed algorithm introduces a modification to the modular range reduction algorithm which increases the speed of computation and allows us to design an architecture for the floating-point case. The implementation presented admits as an input argument any representable number of the standard single precision IEEE 754 floating-point representation and provides the maximum accuracy to the final result. This supposes a hardware solution to the problem of having an input argument close to a multiple of the constant. A final comparison with other implementations is presented.
Julio Villalba, Tomás Lang, Mario A. González
IEEE Trans. Computers1
2005 Low Latency Pipelined Circular CORDIC
abstract
The pipelined CORDIC with linear approximation to rotation has been proposed to achieve reductions in delay, power and area; however, the schemes for rotation (multiplication) and vectoring (division) complicate implementation in a single unit. In this work, we improve the linear approximation scheme, leading to a unified implementation for rotation and vectoring where fully parallel tree multipliers are used instead of the second half of CORDIC iterations. We also combine the linear approximation to rotation with the scale factor compensation so that the compensation is performed concurrently with the rotation process. Comparison with other designs is also provided.
Elisardo Antelo, Julio Villalba
IEEE Symposium on Computer Arithmetic2
2005 On-line Multioperand Addition Based on On-line Full Adders
abstract
In this paper we deal with the online addition of multioperands for conventional, carry save (CS) and/or signed-digit (SD) numbers. We propose an online full adder (olFA) as the key element to design trees of adders to deal with multioperands. We also consider mixed inputs (e.g. SD and CS) and how to obtain the output in any of these representations. We show that for dealing with online multioperands it is more efficient to work with online CS trees based on olFAs than with online SD tree based on online SD adders. Finally a novel olFA-based architecture is proposed to directly deal with SD numbers.
Julio Villalba, Javier Hormigo, Jose M. Prades, Emilio L. Zapata
ASAP1
2004 Minimum Sum of Absolute Differences Implementation in a Single FPGA Device
Joaquín Olivares 0001, Javier Hormigo, Julio Villalba, José Ignacio Benavides Benítez
FPL3
2004 Evaluation of Elementary Functions Using Multimedia Features
abstract
Summary form only given. Most current computers include multimedia features. We use these extensions to compute elementary functions based on polynomial approximations. Hence, we present several alternatives taking advantage of the new attributes on multimedia processors, such as VLIW and SIMD architectures. Our algorithms support the polynomial evaluation in two different ways: the first one is only based in addition/shift operations; while the second uses MAC instructions. Both approximations are analyzed and tailored to subword parallelism units of the new processors. Potential instruction-level and machine-level parallelism are fully exploited through concurrent use of all functional units. A combined approximation using MAC units and addition and shifts is also presented as a third approximation. Two new instructions are also presented here to improve the execution of some of our algorithms.
Gerardo Bandera, Mario A. González, Julio Villalba, Javier Hormigo, Emilio L. Zapata
IPDPS3
2002 Polynomial Evaluation on Multimedia Processors
abstract
In this paper we deal with polynomial evaluation based on new processor architectures for multimedia applications. We introduce some algorithms to take advantage of the new attributes of multimedia processors, such as VLIW (very long instruction word) and SIMD (single instruction multiple data architecture) architectures. Algorithms to support polynomial evaluation based only in addition/shift operations and other different algorithms with MAC (multiply-and-add) instructions are analyzed and tailored to subword parallelism units of the new processors. Both potential instruction-level and machine-level parallelism are fully exploited through concurrent use of all functional units.
Julio Villalba, Gerardo Bandera, Mario A. González, Javier Hormigo, Emilio L. Zapata
ASAP1
2000 A Hardware Algorithm for Variable-Precision Logarithm
abstract
This paper presents an efficient hardware algorithm for variable-precision logarithm. The algorithm uses an iterative technique that employs table lookups and polynomial approximations. Compared to similar algorithms, it reduces the number of fixed-precision operations by avoiding full precision computations and dynamically varying the precision of intermediate results. It also uses significantly smaller tables than related algorithms. For a specified hardware implementation, the algorithm requires fewer than 2L/sup 2/ fixed-precision multiplications to evaluate the logarithm to L words of precision. An error analysis for the algorithm is also presented.
Javier Hormigo, Julio Villalba, Michael J. Schulte
ASAP2
1999 Interval Sine and Cosine Functions Computation Based on Variable-Precision CORDIC Algorithm
abstract
In this paper we design a CORDIC architecture for variable-precision, and a new algorithm is proposed to perform the interval sine and cosine functions. This system allows us to specify the precision to perform the sine and cosine functions, and control the accuracy of the result, in such a way that recomputation of inaccurate results can be carried out with higher precision. An important reduction in the number of iterations is obtained by taking advantage of the differential angle, and the number of cycles per iteration is reduced by avoiding the additions of the leading all zero words. As a consequence, the computation time of the interval function evaluation obtained is close to that of a point function evaluation. The problem of the large table of angles and the scale factor compensation involved in a high precision CORDIC has been solved efficiently.
Javier Hormigo, Julio Villalba, Emilio L. Zapata
IEEE Symposium on Computer Arithmetic2
1997 Low latency word serial CORDIC
abstract
In this paper we present a modification of the CORDIC algorithm which reduces the number of iterations almost to half by merging two successive iterations of the basic algorithm. The two coefficients per iteration are obtained with only a small increase in the cycle time by estimating one of the coefficients. A correcting iteration method is used to correct the possible errors produced by the estimate. Moreover, the modified iteration permits the reduction of the number of cycles required for the compensation of the scaling factor. The resulting architecture is word serial, working both in rotation and vectoring operation modes, presenting a low latency in comparison with the classical CORDIC approach.
Julio Villalba, Tomás Lang
ASAP1
1997 High Performance Rotation Architectures Based on the Radix-4 CORDIC Algorithm
abstract
Traditionally, CORDIC algorithms have employed radix-2 in the first n/2 microrotations (n is the precision in bits) in order to preserve a constant scale factor. The authors present a full radix-4 CORDIC algorithm in rotation mode and circular coordinates and its corresponding selection function, and propose an efficient technique for the compensation of the nonconstant scale factor. Three radix-4 CORDIC architectures are implemented: 1) a word serial architecture based on the zero skipping technique, 2) a pipelined architecture, and 3) an application specific architecture (the angles are known beforehand). The first two are general purpose implementations where redundant (carry-save) or nonredundant arithmetic can be used, whereas the last one is a simplification of the first two. The proposed architectures present a good trade-off between latency and hardware complexity when compared with existing CORDIC architectures.
Elisardo Antelo, Julio Villalba, Javier D. Bruguera, Emilio L. Zapata
IEEE Trans. Computers2
1996 Radix-4 Vectoring Cordic Algorithm And Architectures
abstract
In this paper we present a new CORDIC algorithm for the vectoring mode, based on the use of radix-4 preserving a complexity in the microrotations that is similar to that of the conventional radix-2 CORDIC. The use of this radix, together with the inclusion in the CORDIC algorithm of the zero skipping technique, reduces by more than half the number of iterations with respect to the conventional radix 2 CORDIC, with the consequent reduction of time in recursive architectures or area in pipelined architectures. In processes such as SVD or matrix triangularization in which the evaluation of the rotation angle is required, this algorithm is shown to be specially efficient.
Julio Villalba, J. C. Arrabal, Emilio L. Zapata, Elisardo Antelo, Javier D. Bruguera
ASAP1
1995 Redundant CORDIC Rotator Based on Parallel Prediction
abstract
We present a Cordic rotator, using carry-save arithmetic, based on the prediction of all the coefficients into which the rotation angle is decomposed. The prediction algorithm is based on the use of radix-2 microrotations with multiple shifts in the first iterations and the use of a redundant radix-2 and radix-4 representation for the coefficients in the rest of the microrotations. The use of multiple shifts facilitates the prediction of the coefficients in the case of microrotations where i/spl les/n/4, being n the precision of the algorithm, and the use of radix-4 microrotations helps to reduce the total number of iterations. The prediction is carried out using the redundant representation of the z coordinate, without any need for conversions to a non-redundant representation. Finally, we present a VLSI architecture based on this algorithm. As the production of the coefficients is very fast, and they are known before starting each microrotation, the resulting architecture can be highly pipelined and consequently appropriate for applications where high speeds are required.>
Elisardo Antelo, Javier D. Bruguera, Julio Villalba, Emilio L. Zapata
IEEE Symposium on Computer Arithmetic3
1995 Digit On-line Large Radix CORDIC Rotator
abstract
Many applications figure the evaluation of rotations at high speeds. However there is a trade-off between the chip area and the latency. In this paper we develop a digit on-line pipelined array architecture based on the radix-4 CORDIC algorithm in rotation mode. The radix-4 CORDIC algorithm halves the number of microrotations with respect the traditionally radix-2 algorithm with the drawback of a non-constant scale factor. Seeking a good compromise between silicon area and latency we have used digit on-line processing. This way the data inputs the processor in blocks of bits (digits) in MSD-first mode of processing. We have used redundant carry-save arithmetic to allow carry-free additions and on-line processing. The designed processor demonstrates to have a better performance than previous digit on-line architectures.
Roberto R. Osorio, Elisardo Antelo, Javier D. Bruguera, Julio Villalba, Emilio L. Zapata
ASAP4
1995 CORDIC Architectures with Parallel Compensation of the Scale Factor
abstract
The compensation of scale factor imposes significant computation overhead on the CORDIC algorithm. In this paper we will propose two algorithms and architectures in order to perform the compensation of the scale factor in parallel with the computation of the CORDIC iterations. This way it is not necessary to carry out the final multiplication or add scaling iterations in order to achieve the compensation. With the architectures we propose the dependence on n of the compensation of the scale factor disappears, and this considerably reduces the latency of the system. The architectures developed are optimized solutions for the different operating modes of the CORDIC both in conventional and in redundant arithmetic.
Julio Villalba, José A. Hidalgo-López, Emilio L. Zapata, Elisardo Antelo, Javier D. Bruguera
ASAP1
1995 A fast Hough transform for segment detection
abstract
The authors describe a new algorithm for the fast Hough transform (FHT) that satisfactorily solves the problems other fast algorithms propose in the literature-erroneous solutions, point redundance, scaling, and detection of straight lines of different sizes-and needs less storage space. By using the information generated by the algorithm for the detection of straight lines, they manage to detect the segments of the image without appreciable computational overhead. They also discuss the performance and the parallelization of the algorithm and show its efficiency with some examples.
Nicolás Guil, Julio Villalba, Emilio L. Zapata
IEEE Trans. Image Process.2