Mario Garrido

dblp:79/10311 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-5739-3544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Advanced Quantization Schemes to Increase Accuracy, Reduce Area, and Lower Power Consumption in FFT Architectures
abstract
This paper explores new advanced quantization schemes for fast Fourier transform (FFT) architectures. In previous works, FFT quantization has been treated theoretically or with the sole aim of improving accuracy. In this work, we go one step beyond by considering also the implications that quantization schemes have on the area and power consumption of the architecture. To achieve this, we have analyzed the mathematical operations carried out in FFT architectures and explored the changes that benefit all the figures of merit. By combining or alternating truncation and rounding, and using the half-unit biased (HUB) representation in the different computations of the architecture, we have achieved quantization schemes that increase accuracy, reduce area, and lower power consumption simultaneously. This win-win result improves multiple figures of merit without worsening any other, making it a valuable strategy to optimize FFT architectures.
Mario Garrido, Víctor Manuel Bautista, Alejandro Portas, Javier Hormigo
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Uniformly Distributed CORDIC
abstract
This paper presents the uniformly distributed (UD) coordinate rotation digital computer (CORDIC). It receives this name because the distance between its rotation angles is uniform. The paper presents two versions of the UD CORDIC: the binary UD CORDIC and the canonical signed digit (CSD) UD CORDIC. Both simplify the control part of the CORDIC by obtaining the micro-rotation directions straightforwardly from the binary representation of the rotation angle, with the particularity that, in the CSD UD CORDIC, the binary representation is transformed first into a CSD number. Compared to previous CORDIC approaches that simplify the control part, the proposed rotators further simplify the micro-rotation stages. This is achieved by using more efficient angles for the micro-rotations in the binary UD CORDIC and by merging consecutive micro-rotation stages in the CSD UD CORDIC. Thus, the simplification of the control part of the rotators, the careful selection of the micro-rotation angles, and their optimized shift-and-add implementation results in CORDIC architectures that require the smallest number of adders among pipelined CORDIC implementations for a given resolution so far. Additionally, experimental results of the proposed approach on field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) are provided, showing improvements with respect to previous approaches in multiple figures of merit, mostly area, clock frequency, and power consumption.
Mario Garrido, Daniel Medina, Pedro Paz, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 Novel Access Patterns Based on Overlapping Loading and Processing Times to Reduce Latency and Increase Throughput in Memory-based FFTs
abstract
This paper presents novel access patterns for P-parallel N-point radix-2 memory-based fast Fourier transform (FFT) architectures. This work aims to reduce the latency and increase the throughput by changing the data order and/or choosing different places of the architectures to input/output data. In this way, we can eliminate the loading time and/or the time to collect the output data. This results in a reduction in latency and an increase in throughput. Likewise, the architectures use the same permutation circuits for each iteration, which simplifies the circuit. In addition to improvements in latency and throughput with different access patterns for memory-based FFTs, this work also offers bit-reversed or natural input/output alternatives.
Zeynep Kaya, Mario Garrido
ARITH2
2023 Energy-Efficient Short-Time Fourier Transform for Partial Window Overlapping
abstract
This paper presents an energy-efficient short-time Fourier transform (STFT) architecture. The proposed architecture is called frequency decomposition STFT (FD-STFT) and it achieves significant computational complexity reduction by effectively re-utilizing previously computed spectrums between overlapped sampling windows. Such an algorithmic modification not only reduces the required hardware units, but also achieves low accumulative error compared to conventional approaches. In addition, the quality of the resulting spectrogram is improved by integrating an efficient Hanning windowing technique that replaces the multiplication in the time domain with a low-cost filtering in the frequency domain. For an$N=256$-point window with$R=32$overlapping samples, our results indicate that our approach achieves up-to 40.86% and 65.56% area and power savings respectively, compared to recent approaches.
Charalampos Eleftheriadis, Mario Garrido, Georgios Karakonstantis
ISCAS2
2023 Serial Butterflies for Non-Power-of-Two FFT Architectures in 5G and Beyond
abstract
This paper presents new serial butterflies for non-power-of-two (NP2) fast Fourier transform (FFT) architectures. The paper considers radices 2, 3, 4, and 5, which are used in FFTs for 5G systems. Current designs for non-power-of-two FFTs are mostly based on the single-path delay feedback (SDF) architecture. This type of architecture processes data arriving in series. However, it uses butterflies with several parallel inputs. This results in low utilization, as the butterflies have to wait for all the inputs before they start to process them. Conversely, the proposed approach allows to calculate the butterflies on data that arrive in series. This removes waiting times and reduces the number of hardware components such as multipliers and adders. As a result, the proposed butterflies achieve high performance and provide a significant reduction in area and power consumption with respect to parallel butterflies. Thus, they are an efficient solution when data must be processed in series in the butterflies.
Víctor Manuel Bautista, Mario Garrido, Marisa López-Vallejo
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Low-Latency 64-Parallel 4096-Point Memory-Based FFT for 6G
abstract
This paper presents a novel 64-parallel 4096-point radix-2 memory-based fast Fourier transform (FFT) architecture for 6G. This approach is the first one to use 64 parallel branches in memory-based architectures. The challenge of designing a memory-based FFT with such a high parallelization has been accomplished by paying special attention to the large number of memories in parallel. Their control has been simplified by using the same read and write address for all of them thanks to the perfect shuffle permutation, and they are organized in groups to eliminate unnecessary registers. Likewise, a novel design for the rotation memories allows for reusing rotation coefficients among parallel rotators, and a new design for the circular counter that controls the architecture is presented. The proposed FFT architecture has been implemented on a Virtex 7 field-programmable gate array (FPGA). Experimental results reveal that the proposed architecture achieves the lowest latency in clock cycles and the highest throughput in samples per clock cycle among memory-based FFT architectures so far.
Zeynep Kaya, Mario Garrido
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 The Constant Multiplier FFT
abstract
In this paper, we present a new fast Fourier transform (FFT) hardware architecture called constant multiplier (CM) FFT. Whereas rotators in previous architectures must rotate among several different angles, the CM FFT exploits the use of constant multipliers to calculate the rotations. The paper explores the 4-parallel and 8-parallel radix-2 CM FFT, which calculates the FFT entirely by using constant multipliers. Later, radices 2kare presented, with emphasis in radix-24and radix-25as the best alternatives. Experimental results for a 1024-point radix-25CM FFT show that the proposed approach reduces the number of block random-access memories (BRAM) and digital signal processing (DSP) slices with respect to previous approaches, while achieving the highest clock frequency for a 4-parallel FFT architecture on field-programmable gate arrays (FPGAs) reported so far.
Mario Garrido, Pedro Malagón
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 A 128-Point Multi-Path SC FFT Architecture
abstract
This paper presents a new radix-2kmulti-path FFT architecture, named MSC FFT, which is based on a single-path radix-2 serial commutator (SC) FFT architecture. The proposed multi-path architecture has a very high hardware utilization that results in a small chip area, while providing high throughput. In addition, the adoption of radix-2kFFT algorithms allows for simplifying the rotators even further. It is achieved by optimizing the structure of the processing element (PE). The implemented architecture is a 128-point 4-parallel multi-path SC FFT using 90 nm process. Its area and power consumption at 250 MHz are only 0.167 mm2and 14.81 mW, respectively. Compared with existing works, the proposed design reduces significantly the chip area and the power consumption, while providing high throughput.
Shun-Che Hsu, Shen-Jui Huang, Sau-Gee Chen, Shin-Che Lin, Mario Garrido
ISCAS5
2019 Effect of Finite Word-Length on SQNR, Area and Power for Real-Valued Serial FFT
abstract
Modern applications for DSP systems are increasingly constrained by tight area and power requirements. Therefore, it is imperative to analyze effective strategies that work within these requirements. This paper studies the impact of finite word-length arithmetic on the signal to quantization noise ratio (SQNR), power and area for a real-valued serial FFT implementation. An experiment is set up using a hardware description language (HDL) to empirically determine the tradeoffs associated with the following parameters: (i) the input word-length, (ii) the word-length of the rotation coefficients, and (iii) length of the FFT on performance (SQNR), power and area. The results of this paper can be used to make design decisions by careful selection of word-length to achieve a reduction in area and power for an acceptable loss in SQNR.
Nanda K. Unnikrishnan, Mario Garrido, Keshab K. Parhi
ISCAS2
2019 Optimum Circuits for Bit-Dimension Permutations
abstract
In this paper, we present a systematic approach to design hardware circuits for bit-dimension permutations. The proposed approach is based on decomposing any bit-dimension permutation into elementary bit-exchanges. Such decomposition is proven to achieve the theoretical minimum number of delays required for the permutation. This offers optimum solutions for multiple well-known problems in the literature that make use of bit-dimension permutations. This includes the design of permutation circuits for the fast Fourier transform, bit reversal, matrix transposition, stride permutations, and Viterbi decoders.
Mario Garrido, Jesús Grajal, Oscar Gustafsson
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Optimal Shift Reassignment in Reconfigurable Constant Multiplication Circuits
abstract
This paper presents a new method called optimal shift reassignment (OSR), used for reconfigurable multiplication circuits. These circuits consist of adders, subtractors, shifts, and multiplexers (MUXs). They calculate the multiplication of an input number by one out of several constants which can be selected dynamically during run-time. The OSR method is based on the idea that shifts can be placed at different positions along the circuit, while the calculated output constant stays the same. This differs from previous approaches, which were limited by the fact that all constants within the constant multiplier were forced to be odd. The OSR method subsequently releases this restriction. As a result, the number of required MUXs in the circuit can be reduced. This happens when the shift reassignment aligns the shift values of different inputs of an MUX. Experimental results show MUX savings of up to 50% and average savings between 11% and 16% using the OSR method compared to previous approaches.
Konrad Möller, Martin Kumm, Mario Garrido, Peter Zipf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 A 4096-Point Radix-4 Memory-Based FFT Using DSP Slices
abstract
This brief presents a novel 4096-point radix-4 memory-based fast Fourier transform (FFT). The proposed architecture follows a conflict-free strategy that only requires a total memory of size N and a few additional multiplexers. The control is also simple, as it is generated directly from the bits of a counter. Apart from the low complexity, the FFT has been implemented on a Virtex-5 field programmable gate array (FPGA) using DSP slices. The goal has been to reduce the use of distributed logic, which is scarce in the target FPGA. With this purpose, most of the hardware has been implemented in DSP48E. As a result, the proposed FPGA is efficient in terms of hardware resources, as is shown by the experimental results.
Mario Garrido, Miguel Angel Sánchez, Marisa López-Vallejo, Jesús Grajal
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Multiplierless Unity-Gain SDF FFTs
abstract
In this brief, we propose a novel approach to implement multiplierless unity-gain single-delay feedback fast Fourier transforms (FFTs). Previous methods achieve unity-gain FFTs by using either complex multipliers or nonunity-gain rotators with additional scaling compensation. Conversely, this brief proposes unity-gain FFTs without compensation circuits, even when using nonunity-gain rotators. This is achieved by a joint design of rotators, so that the entire FFT is scaled by a power of two, which is then shifted to unity. This reduces the amount of hardware resources of the FFT architecture, while having high accuracy in the calculations. The proposed approach can be applied to any FFT size, and various designs for different FFT sizes are presented.
Mario Garrido, Rikard Andersson, Fahad Qureshi, Oscar Gustafsson
IEEE Trans. Very Large Scale Integr. Syst.1
2013 A reconfigurable FFT architecture for variable-length and multi-streaming OFDM standards
abstract
This paper presents a reconfigurable FFT architecture for variable-length and multi-streaming WiMax wireless standard. The architecture processes 1 stream of 2048-point FFT, up to 2 streams of 1024-point FFT or up to 4 streams of 512-point FFT. The architecture consists of a modified radix-2 single delay feedback (SDF) FFT. The sampling frequency of the system is varied in accordance with the FFT length. The latch-free clock gating technique is used to reduce power consumption. The proposed architecture has been synthesized for the Virtex-6 XCVLX760 FPGA. Experimental results show that the architecture achieves the throughput that is required by the WiMax standard and the design has additional features compared to the previous approaches. The design uses 1% of the total available FPGA resources and maximum clock frequency of 313.67 MHz is achieved. Furthermore, this architecture can be expanded to suit other wireless standards.
Padma Prasad Boopal, Mario Garrido, Oscar Gustafsson
ISCAS2
2013 Pipelined Radix-2k Feedforward FFT Architectures
abstract
The appearance of radix-22was a milestone in the design of pipelined FFT hardware architectures. Later, radix-22was extended to radix-2k. However, radix-2kwas only proposed for single-path delay feedback (SDF) architectures, but not for feedforward ones, also called multi-path delay commutator (MDC). This paper presents the radix-2kfeedforward (MDC) FFT architectures. In feedforward architectures radix-2kcan be used for any number of parallel samples which is a power of two. Furthermore, both decimation in frequency (DIF) and decimation in time (DIT) decompositions can be used. In addition to this, the designs can achieve very high throughputs, which makes them suitable for the most demanding applications. Indeed, the proposed radix-2kfeedforward architectures require fewer hardware resources than parallel feedback ones, also called multi-path delay feedback (MDF), when several samples in parallel must be processed. As a result, the proposed radix-2kfeedforward architectures not only offer an attractive solution for current applications, but also open up a new research line on feedforward structures.
Mario Garrido, Jesús Grajal, Miguel A. Sánchez Marcos, Oscar Gustafsson
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Efficient Memoryless Cordic for FFT Computation
abstract
A new memoryless CORDIC algorithm for the FFT computation is proposed in this paper. This approach calculates the direction of the micro-rotations from the control counter of the FFT, so the area of the rotator hardly depends on the number of rotations, which is particularly suitable for the computation of FFTs of a high number of points. Moreover, the new CORDIC presents other advantages such as the simplification of the basic CORDIC processor used to calculate the micro-rotations, or an easy way to compensate the intrinsic gain of the CORDIC algorithm. Additionally, the VLSI implementation of the algorithm is a pipeline architecture with high performance in terms of speed, throughput and latency.
Mario Garrido, Jesús Grajal
ICASSP (2)1
2006 Automated design space exploration of FPGA-based FFT architectures based on area and power estimation
abstract
In this paper a tool aimed at generating fast Fourier transform (FFT) cores targeting FPGA platforms was presented. The tool is able to generate different pipelined architectures of the FFT that provide different points of the design space: from high performance to low area implementations. The user can select the most suitable architecture based on a broad set of configuration parameters, as they are the number of points, sample size, truncation, etc. Moreover, a set of accurate estimators has been implemented to allow the designer an early and quick design space exploration before synthesizing the core. Experimental results validate our approach and provide significant measurements about the accuracy of the estimation and the tool execution time
Miguel A. Sánchez Marcos, Mario Garrido, Marisa López-Vallejo, Carlos A. López-Barrio
FPT2