VLDB 2026 Research / reviewers in the wild / expert
Javier Hormigo
dblp:58/7026
· DBLP profile ↗
31ranked-venue papers
10as first author
3since 2021 · last 2025
0000-0002-5454-6821ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 7 first-author · 2 since 2021Theory of computation · 6 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging SYCL for Heterogeneous cDTW Computation on CPU, GPU, and FPGAabstractABSTRACT One of the most time‐consuming kernels of a recent epileptic seizure detection application is the computation of the constrained Dynamic Time Warping (cDTW) Distance Matrix. In this paper, we explore the design space of heterogeneous CPU, GPU, and FPGA implementations of this kernel using SYCL as a programming model. First, we optimize the CPU implementation leveraging the SIMD capability of SYCL and compare it with the latest C++26 SIMD library. Next, we tune the SYCL code to run on an on‐chip GPU, iGPU, as well as on a discrete NVIDIA GPU, dGPU. We also develop a SYCL implementation on an Intel FPGA. On top of that, we exploit simultaneous co‐processing on CPU+GPU and CPU+FPGA platforms by extending a previous heterogeneous scheduling framework to now support 2D partitioning strategies. Our evaluations demonstrate that SYCL seems well suited to exploit the SIMD capabilities of modern CPU cores and shows promising results for accelerating devices, both in terms of performance and energy efficiency. Moreover, we find that our scheduler enables the efficient co‐execution of work among the computing devices, and the results demonstrate that dynamic and adaptive partitioning strategies perform efficiently with overheads below 4%. Cristian Campos, Rafael Asenjo, Javier Hormigo, Angeles G. Navarro |
Concurr. Comput. Pract. Exp. | 3 |
| 2025 | Advanced Quantization Schemes to Increase Accuracy, Reduce Area, and Lower Power Consumption in FFT ArchitecturesabstractThis paper explores new advanced quantization schemes for fast Fourier transform (FFT) architectures. In previous works, FFT quantization has been treated theoretically or with the sole aim of improving accuracy. In this work, we go one step beyond by considering also the implications that quantization schemes have on the area and power consumption of the architecture. To achieve this, we have analyzed the mathematical operations carried out in FFT architectures and explored the changes that benefit all the figures of merit. By combining or alternating truncation and rounding, and using the half-unit biased (HUB) representation in the different computations of the architecture, we have achieved quantization schemes that increase accuracy, reduce area, and lower power consumption simultaneously. This win-win result improves multiple figures of merit without worsening any other, making it a valuable strategy to optimize FFT architectures. Mario Garrido, Víctor Manuel Bautista, Alejandro Portas, Javier Hormigo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | FPGA acceleration of bit-true simulations for word-length optimizationabstractThe end of Moore's law and the arrival of new highly demanding applications have awakened the interest in exploring different number representation formats and also combining them to implement domain-specific accelerators. Typically used in DSP applications, word-length optimization (WLO) allows finding the optimum combination of word-lengths for each signal on a circuit for a given error threshold. In the optimization process, for any word-length combination, the error has to be estimated or computed by bit-true simulation. The latter is widely used since it can be applied to any type of system. However, simulation is very time-consuming, and the WLO becomes an extremely long process. This paper proposes a methodology based on a WLO-wise hardware architecture that speeds up WLO significantly. In our approach, the target datapath is implemented on an FPGA with a “precision limiter” on each selected signal. This architecture allows performing bit-true emulation on the FPGA for any given word-length combination without reconfiguring the FPGA; just configuring the limiters, which is a much faster process. Javier Hormigo, Gabriel Caffarena |
ARITH | 1 |
| 2020 | Floating-Point Fused Multiply-Add under HUB FormatabstractThe Half-Unit-Biased (HUB) format has interesting advantages for implementing floating-point arithmetic which has been proved for the four basic arithmetic operations as well as square root. Nevertheless, although Floating-point Fused Multiply-add (FMA) operation (AxB + C) is one of the most important and complex arithmetic instructions in modern processors, FMA operation for HUB numbers has not been confronted yet. In this paper, we present a design to deal with this operation under HUB format. The key points to turn the conventional FMA architecture into a HUB unit are explained. Comparing the ASIC implementation of a HUB FMA unit with the conventional one, the former reduces the required area and power up to 38% and 35%, respectively, for single-precision. For BFloat16, the HUB FMA increases the speed a 15%, and even then, reduces the area and power by 26% and 12%, respectively. Javier Hormigo, Julio Villalba, Sonia Gonzalez-Navarro |
ARITH | 1 |
| 2020 | New Results on Non-Normalized Floating-Point FormatsabstractCompulsory normalization of the represented numbers is a key requirement of the floating-point standard. This requirement contributes to fundamental characteristics of the standard, such as taking the most of the precision available, reproducibility and facilitation of comparison and other operations. However, it also imposes a high restriction in effectiveness of basic arithmetic operation implementation. In many embedded applications may be worth to sacrifice the benefits of normalization for gaining in implementation metrics. This paper analyzes and measures the effect of removing the normalization requirement in terms of precision and implementation savings for embedded applications. We propose several adder and multiplier architectures to deal with non-normalized floating-point numbers, and quantify the accuracy loss and the improvements in hardware implementation. Our experiments show that it is possible to reduce the area and power consumption up to 78 percent in ASIC and 50 percent in FPGA implementations with a reasonable accuracy loss. Sonia Gonzalez-Navarro, Javier Hormigo |
IEEE Trans. Computers | 2 |
| 2019 | Reproducible Summation Under HUB FormatabstractFloating point reproducibility is a property claimed by programmers and end users. Half-Unit-Biased (HUB) is a new representation format in which the round to nearest is carried out by truncation, preventing any carry propagation and saving time and area. In this paper we study the reproducible summation of HUB numbers by using a error-free vector transformation technique, providing both a specific architecture and the usage of combined HUB/Standard floating point adders to achieve a reproducible result. Julio Villalba, Javier Hormigo, Francisco J. Jaime |
ARITH | 2 |
| 2018 | Unbiased Rounding for HUB Floating-Point AdditionabstractHalf-Unit-Biased (HUB) is an emerging format based on shifting the represented numbers by half Unit in the Last Place. This format simplifies two's complement and round-to-nearest operations by preventing any carry propagation. This saves power consumption, time and area. Taking into account that the IEEE floating-point standard uses an unbiased rounding as the default mode, this feature is also desirable for HUB approaches. In this paper, we study the unbiased rounding for HUB floating-point addition in both as standalone operation and within FMA. We show two different alternatives to eliminate the bias when rounding the sum results, either partially or totally. We also present an error analysis and the implementation results of the proposed architectures to help the designers to decide what their best option are. Julio Villalba, Javier Hormigo, Sonia Gonzalez-Navarro |
IEEE Trans. Computers | 2 |
| 2017 | Normalizing or Not Normalizing? An Open Question for Floating-Point Arithmetic in Embedded SystemsabstractEmerging embedded applications lack of a specific standard when they require floating-point arithmetic. In this situation they use the IEEE-754 standard or ad hoc variations of it. However, this standard was not designed for this purpose. This paper aims to open a debate to define a new extension of the standard to cover embedded applications. In this work, we only focus on the impact of not performing normalization. We show how eliminating the condition of normalized numbers, implementation costs can be dramatically reduced, at the expense of a moderate loss of accuracy. Several architectures to implement addition and multiplication for non-normalized numbers are proposed and analyzed. We show that a combined architecture (adder-multiplier) can halve the area and power consumption of its counterpart IEEE-754 architecture. This saving comes at the cost of reducing an average of about 10 dBs the Signal-to-Noise Ratio for the tested algorithms. We think these results should encourage researchers to perform further investigation in this issue. Sonia Gonzalez-Navarro, Javier Hormigo |
ARITH | 2 |
| 2017 | Floating Point Square Root under HUB FormatabstractUnit-Biased (HUB) is an emerging format based on shifting the representation line of the binary numbers by half unit in the last place. The HUB format is specially relevant for computers where rounding to nearest is required because it is performed simply by truncation. From a hardware point of view, the circuits implementing this representation save both area and time since rounding does not involve any carry propagation. Designs to perform the four basic operations have been proposed under HUB format recently. Nevertheless, the square root operation has not been confronted yet. In this paper we present an architecture to carry out the square root operation under HUB format for floating point numbers. The results of this work keep supporting the fact that the HUB representation involves simpler hardware than its conventional counterpart for computers requiring round-to-nearest mode. Julio Villalba, Javier Hormigo |
ICCD | 2 |
| 2017 | Introduction to the Special Issue on Computer ArithmeticabstractThe papers in this special issue focus on computer arithmetic which is used in many applications, usually totally silently (one should keep in mind that even when running programs that are not at all numeric, memory addresses are computed, which involves additions, multiplications, and sometimes divisions). However, in some areas, it plays a central role. Javier Hormigo, Jean-Michel Muller, Stuart F. Oberman, Nathalie Revol, Arnaud Tisserand, Julio Villalba |
IEEE Trans. Computers | 1 |
| 2016 | New Formats for Computing with Real-Numbers under Round-to-NearestabstractIn this paper, a new family of formats to deal with real number for applications requiring round to nearest is proposed. They are based on shifting the set of exactly represented numbers which are used in conventional radix-$\beta$number systems. This technique allows performing radix complement and round to nearest without carry propagation with negligible time and hardware cost. Furthermore, the proposed formats have the same storage cost and precision as standard ones. Since conversion to conventional formats simply require appending one extra-digit to the operands, standard circuits may be used to perform arithmetic operations with operands under the new format. We also extend the features of the RN-representation system and carry out a thorough comparison between both representation systems. We conclude that the proposed representation system is generally more adequate to implement systems for computation with real number under round-to-nearest. Javier Hormigo, Julio Villalba |
IEEE Trans. Computers | 1 |
| 2016 | Measuring Improvement When Using HUB Formats to Implement Floating-Point Systems Under Round-to-NearestabstractThis paper analyzes the benefits of using half-unit-biased (HUB) formats to implement floating-point (FP) arithmetic under a round-to-nearest mode from a quantitative point of view. Using the HUB formats to represent numbers allows the removal of the rounding logic of arithmetic units, including sticky-bit computation. This is shown for FP adders, multipliers, and converters. Experimental analysis demonstrates that the HUB formats and the corresponding arithmetic units maintain the same accuracy as the conventional ones. On the other hand, the implementation of these units, based on basic architectures, shows that the HUB formats simultaneously improve area, speed, and power consumption. In addition, based on the data obtained from the synthesis, an HUB single-precision adder is ~14% faster but consumes 38% less area and 26% less power than the conventional adder. Similarly, an HUB single-precision multiplier is 17% faster, uses 22% less area, and consumes slightly less power than the conventional multiplier. At the same speed, the adder and the multiplier achieve area and power reductions of up to 50% and 40%, respectively. Javier Hormigo, Julio Villalba |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Efficient floating-point representation for balanced codes for FPGA devicesabstractWe propose a floating-point representation to deal efficiently with arithmetic operations in codes with a balanced number of additions and multiplications for FPGA devices. The variable shift operation is very slow in these devices. We propose a format that reduces the variable shifter penalty. It is based on a radix-64 representation such that the number of the possible shifts is considerably reduced. Thus, the execution time of the floating-point addition is highly optimized when it is performed in an FPGA device, which compensates for the multiplication penalty when a high radix is used, as experimental results have shown. Consequently, the main problem of previous specific high-radix FPGA designs (no speedup for codes with a balanced number of multiplications and additions) is overcome with our proposal. The inherent architecture supporting the new format works with greater bit precision than the corresponding single precision (SP) IEEE-754 standard. Julio Villalba, Javier Hormigo, Francisco Corbera, Mario A. González, Emilio L. Zapata |
ICCD | 2 |
| 2013 | Multioperand Redundant Adders on FPGAsabstractAlthough redundant addition is widely used to design parallel multioperand adders for ASIC implementations, the use of redundant adders on Field Programmable Gate Arrays (FPGAs) has generally been avoided. The main reasons are the efficient implementation of carry propagate adders (CPAs) on these devices (due to their specialized carry-chain resources) as well as the area overhead of the redundant adders when they are implemented on FPGAs. This paper presents different approaches to the efficient implementation of generic carry-save compressor trees on FPGAs. They present a fast critical path, independent of bit width, with practically no area overhead compared to CPA trees. Along with the classic carry-save compressor tree, we present a novel linear array structure, which efficiently uses the fast carry-chain resources. This approach is defined in a parameterizable HDL code based on CPAs, which makes it compatible with any FPGA family or vendor. A detailed study is provided for a wide range of bit widths and large number of operands. Compared to binary and ternary CPA trees, speedups of up to 2.29 and 2.14 are achieved for 16-bit width and up to 3.81 and 3.11 for 64-bit width. Javier Hormigo, Julio Villalba, Emilio L. Zapata |
IEEE Trans. Computers | 1 |
| 2013 | Self-Reconfigurable Constant Multiplier for FPGAabstractConstant multipliers are widely used in signal processing applications to implement the multiplication of signals by a constant coefficient. However, in some applications, this coefficient remains invariable only during an interval of time, and then, its value changes to adapt to new circumstances. In this article, we present a self-reconfigurable constant multiplier suitable for LUT-based FPGAs able to reload the constant in runtime. The pipelined architecture presented is easily scalable to any multiplicand and constant sizes, for unsigned and signed representations. It can be reprogrammed in 16 clock cycles, equivalent to less than 100 ns in current FPGAs. This value is significantly smaller than FPGA partial configuration times. The presented approach is more efficient in terms of area and speed when compared to generic multipliers, achieving up to 91% area reduction and up to 102% speed improvement for the case-study circuits tested. The power consumption of the proposed multipliers are in the range of those of slice-based multipliers provided by the vendor. Javier Hormigo, Gabriel Caffarena, Juan P. Oliver, Eduardo I. Boemo |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2012 | A study of decimal left shifters for binary numbers
Sonia Gonzalez-Navarro, Javier Hormigo, Michael J. Schulte |
Inf. Comput. | 2 |
| 2012 | Radix-2 Multioperand and Multiformat Streaming Online AdditionabstractIn this paper, we present multioperand radix-2 online addition using different data representations (signed-digit, two's complement, and carry-save), in particular cases in which operands with different representations are added. We use the previously defined online full adder (olFA) as a component to build different multioperand online architectures. To merge data with different representations, an inner conversion of data is performed, eliminating any conversion stage and penalty time. We propose a technique to build multioperand trees efficiently and give six practical rules to deal with different kinds of data in the same adder. For addition of a stream of data, we determine the minimum number of separation cycles required to isolate two successive computations and propose a novel hardware technique that eliminates completely the separation cycles, resulting in the maximum throughput possible. Julio Villalba, Tomás Lang, Javier Hormigo |
IEEE Trans. Computers | 3 |
| 2011 | High-Speed Algorithms and Architectures for Range Reduction ComputationabstractRange reduction is a crucial step for accuracy in trigonometric functions evaluation. This paper shows and compares a set of algorithms for additive range reduction computation and their corresponding application-specific integrated circuit implementations (ensuring an accuracy of one unit in the last place). A word-serial architecture implementation has been used as a reference for clearer comparisons. Besides, a new table-based pipelined architecture for range reduction has also been proposed. Francisco J. Jaime, M. A. Sánchez, Javier Hormigo, Julio Villalba, Emilio L. Zapata |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Efficient Implementation of Carry-Save Adders in FPGAsabstractMost field programmable gate array (FPGA) devices have a special fast carry propagation logic intended to optimize addition operations. The redundant adders do not easily fit into this specialized carry-logic and, consequently, they require double hardware resources than carry propagate adders, while showing a similar delay for small size operands. Therefore, carry-save adders are not usually implemented on FPGA devices, although they are very useful in ASIC implementations. In this paper we study efficient implementations of carry-save adders on FPGA devices, taking advantage of the specialized carry-logic. We show that it is possible to implement redundant adders with a hardware cost close to that of a carry propagate adder. Specifically, for 16 bits and bigger wordlengths, redundant adders are clearly faster and have an area requirement similar to carry propagate adders. Among all the redundant adders studied, the 4:2 compressor is the fastest one, presents the best exploitation of the logic resources within FPGA slices and the easiest way to adapt classical algorithms to efficiently fit FPGA resources. Javier Hormigo, Manuel Ortiz, Francisco Javier Quiles-Latorre, Francisco J. Jaime, Julio Villalba, Emilio L. Zapata |
ASAP | 1 |
| 2008 | SIMD Enhancements for a Hough Transform ImplementationabstractThe Hough transform is a line detection algorithm widely used within image processing applications, showing several variations depending on the shape which is intended to be detected. This paper provides for some new SIMD instructions aided by a specialized look-up table specifically devised for an improved Hough transform algorithm implementation, although they may also be used within other algorithms implementation with a slight modification. The new approach has been tested and evaluated using SimpleScalar tool set. Francisco J. Jaime, Javier Hormigo, Julio Villalba, Emilio L. Zapata |
DSD | 2 |
| 2008 | New SIMD instructions set for image processing applications enhancementabstractDue to its inherent data parallelism, image processing applications benefit from multimedia extensions SIMD instructions within general purpose processors. However, current multimedia extensions do not allow to simultaneously address different memory positions using indirect addressing, i. e. the desired memory position address is located within a register. This restriction forces to sequential execution at some program points. This paper shows a new set of instructions providing parallel indirect addressing to a specialized table and intended to be added to existing multimedia extensions. In order to evaluate the new instructions usefulness and feasibility, we have used SimpleScalar for testing some image processing applications, getting in some cases a speed up of 3.5. Francisco J. Jaime, Javier Hormigo, Julio Villalba, Emilio L. Zapata |
ICIP | 2 |
| 2007 | Improving the Throughput of On-line Addition for Data StreamsabstractIn this paper we deal with the throughput of on–line addition for a stream of data. This throughput is directly related to the initiation interval between two successive instances. The on–line delay for the addition of two signed–digit (or carry–save) numbers is two, and N+2 cycles are classically used to compute a new pair of N–digit data (initiation interval: N+2). In this paper we present some techniques to reduce the initiation interval to N (which is the theoretical minimum value) with a very small amount of hardware or N+1 with no hardware cost. For short operands, this might have a significant effect on the throughput. Julio Villalba, Javier Hormigo, Tomás Lang |
ASAP | 2 |
| 2006 | Pipelined Range Reduction for Floating Point NumbersabstractThis paper presents a new pipelined architecture to deal with range reduction for floating point representation. It is based on Horner's scheme and a look-up table. The overall design has been optimized for a module equal to 2π, which is the most widely used due to trigonometric functions requirements. To ensure an accuracy of one unit in the last place (ULP), a complete error propagation study has been carried out. Francisco J. Jaime, Julio Villalba, Javier Hormigo, Emilio L. Zapata |
ASAP | 3 |
| 2006 | Fast Full-Search Block Matching Algorithm Motion Estimation Alternatives in FPGAabstractBlock matching motion estimation takes a great part of the processing time for video encoding. To accelerate this process is must to reach real time video coding. The best motion vector is obtained by full-search block matching algorithm which has to be usually implemented by hardware. In recent years, several FPGA based designs have been proposed since these devices support high number of process elements in parallel mode. In this paper a survey of recent architectures to perform the full-search block matching algorithm in FPGAs is presented. A further comparison on terms of frames per second reached, hardware cost in CLB slices and system frequency is presented Joaquín Olivares 0001, José Ignacio Benavides Benítez, Javier Hormigo, Julio Villalba, Emilio L. Zapata |
FPL | 3 |
| 2005 | On-line Multioperand Addition Based on On-line Full AddersabstractIn this paper we deal with the online addition of multioperands for conventional, carry save (CS) and/or signed-digit (SD) numbers. We propose an online full adder (olFA) as the key element to design trees of adders to deal with multioperands. We also consider mixed inputs (e.g. SD and CS) and how to obtain the output in any of these representations. We show that for dealing with online multioperands it is more efficient to work with online CS trees based on olFAs than with online SD tree based on online SD adders. Finally a novel olFA-based architecture is proposed to directly deal with SD numbers. Julio Villalba, Javier Hormigo, Jose M. Prades, Emilio L. Zapata |
ASAP | 2 |
| 2004 | Minimum Sum of Absolute Differences Implementation in a Single FPGA Device
Joaquín Olivares 0001, Javier Hormigo, Julio Villalba, José Ignacio Benavides Benítez |
FPL | 2 |
| 2004 | Evaluation of Elementary Functions Using Multimedia FeaturesabstractSummary form only given. Most current computers include multimedia features. We use these extensions to compute elementary functions based on polynomial approximations. Hence, we present several alternatives taking advantage of the new attributes on multimedia processors, such as VLIW and SIMD architectures. Our algorithms support the polynomial evaluation in two different ways: the first one is only based in addition/shift operations; while the second uses MAC instructions. Both approximations are analyzed and tailored to subword parallelism units of the new processors. Potential instruction-level and machine-level parallelism are fully exploited through concurrent use of all functional units. A combined approximation using MAC units and addition and shifts is also presented as a third approximation. Two new instructions are also presented here to improve the execution of some of our algorithms. Gerardo Bandera, Mario A. González, Julio Villalba, Javier Hormigo, Emilio L. Zapata |
IPDPS | 4 |
| 2002 | Polynomial Evaluation on Multimedia ProcessorsabstractIn this paper we deal with polynomial evaluation based on new processor architectures for multimedia applications. We introduce some algorithms to take advantage of the new attributes of multimedia processors, such as VLIW (very long instruction word) and SIMD (single instruction multiple data architecture) architectures. Algorithms to support polynomial evaluation based only in addition/shift operations and other different algorithms with MAC (multiply-and-add) instructions are analyzed and tailored to subword parallelism units of the new processors. Both potential instruction-level and machine-level parallelism are fully exploited through concurrent use of all functional units. Julio Villalba, Gerardo Bandera, Mario A. González, Javier Hormigo, Emilio L. Zapata |
ASAP | 4 |
| 2000 | A Hardware Algorithm for Variable-Precision LogarithmabstractThis paper presents an efficient hardware algorithm for variable-precision logarithm. The algorithm uses an iterative technique that employs table lookups and polynomial approximations. Compared to similar algorithms, it reduces the number of fixed-precision operations by avoiding full precision computations and dynamically varying the precision of intermediate results. It also uses significantly smaller tables than related algorithms. For a specified hardware implementation, the algorithm requires fewer than 2L/sup 2/ fixed-precision multiplications to evaluate the logarithm to L words of precision. An error analysis for the algorithm is also presented. Javier Hormigo, Julio Villalba, Michael J. Schulte |
ASAP | 1 |
| 1999 | Interval Sine and Cosine Functions Computation Based on Variable-Precision CORDIC AlgorithmabstractIn this paper we design a CORDIC architecture for variable-precision, and a new algorithm is proposed to perform the interval sine and cosine functions. This system allows us to specify the precision to perform the sine and cosine functions, and control the accuracy of the result, in such a way that recomputation of inaccurate results can be carried out with higher precision. An important reduction in the number of iterations is obtained by taking advantage of the differential angle, and the number of cycles per iteration is reduced by avoiding the additions of the leading all zero words. As a consequence, the computation time of the interval function evaluation obtained is close to that of a point function evaluation. The problem of the large table of angles and the scale factor compensation involved in a high precision CORDIC has been solved efficiently. Javier Hormigo, Julio Villalba, Emilio L. Zapata |
IEEE Symposium on Computer Arithmetic | 1 |
| 1999 | Texture segmentation through eigen-analysis of the Pseudo-Wigner distribution
Gabriel Cristóbal, Javier Hormigo |
Pattern Recognit. Lett. | 2 |