VLDB 2026 Research / reviewers in the wild / expert
Oscar Gustafsson
dblp:70/3849
· DBLP profile ↗
43ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0003-3470-3911ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 5 first-author · 4 since 2021Theory of computation · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spade - A Modern Hardware Description LanguageabstractThe need for custom hardware to meet compute demands is ever-increasing. However, the hardware description languages (HDLs) that these accelerators are primarily built with were designed in the 1980s which means that they are missing out on over 35 years of development in programming language design. In this article, we present Spade—a new HDL that takes inspiration from modern software language design and aims to be more productive than traditional HDLs without sacrificing performance. This is achieved through a mix of building abstractions for common hardware constructs such as pipelines, and outright borrowing ideas from software, such as a type system that comes close in power to that of Rust or Haskell. Compared to contemporary HDLs such as Chisel, which are embedded in a host language, Spade is a standalone language with a type system that is available in hardware, not just at elaboration-time. Its abstractions also build on top of the RTL abstraction instead of replacing it as is done in languages like BlueSpec and in high-level synthesis. Frans Skarman, Gustav Sörnäs, Oscar Gustafsson |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | Surfer - An Extensible Waveform ViewerabstractAbstract The waveform viewer is one of the most important tools in a hardware engineer’s toolbox. It is the main interface used to track down design bugs found by simulation or formal verification. In this paper, we present Surfer, a modern waveform viewer designed to integrate with the broader hardware design ecosystem. It supports translation from bit vectors to semantically meaningful values, integration with simulation and verification tools, and lays the groundwork for interactive simulation in the open-source ecosystem. Frans Skarman, Lucas Klemmer, Daniel Große, Oscar Gustafsson, Kevin Laeufer |
CAV (4) | 4 |
| 2025 | Low-Complexity Implementation of Real-Time Reconfigurable Low-Pass EqualizersabstractImplementation techniques and results for a recently proposed real-time reconfigurable low-pass equalizer (RLPE) consisting of a variable bandwidth (VBW) filter and a variable equalizer (VE) are presented. Both components utilize fixed finite-length impulse response (FIR) filters combined with a few general multipliers, resulting in lower area and power consumption compared to a general FIR filter, despite requiring more multiplications. This is because the constant multipliers in the fixed FIR filters of the RLPE can be optimized for implementation. An additional advantage is that the proposed RLPE does not require online design. Various implementation alternatives for fixed FIR filters, including ways to increase the frequency, are evaluated to optimize the implementation of the RLPE. Several versions of the proposed RLPE and a general FIR filter for comparison are implemented using a 28-nm fully depleted silicon on insulator (FD-SOI) standard cell library. The results demonstrate that the RLPE baseline design requires less power and area than the general equalizer, and although the frequency of the baseline implementation is lower, the design can reach the same frequency while still having significantly less power and area. Furthermore, an approach is introduced to break the chain in the polynomial section of the VBW filter by using fewer additional registers compared to standard pipelining. Instead, this method reformulates the constant multiplication problem to produce correct results. For the considered case, the power consumption is reduced between 49% and 70% for different frequencies, with an area decrease in the range of 64%–67%, by using the proposed RLPE compared to a general FIR filter. Narges Mohammadi Sarband, Oksana Moryakova, Håkan Johansson, Oscar Gustafsson |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | APyTypes: Algorithmic Data Types in Python for Efficient Simulation of Finite Word-Length EffectsabstractA new Python library, APyTypes, suitable for simulating and exploring finite word-length effects is presented. The library supports configurable bit-accurate fixed- and floatingpoint types of both scalars and multidimensional arrays and uses a C++ backend to accelerate runtime performance. The underlying design principles of the library are introduced and examples show how it can be used. We argue that APyTypes have significant advantages over existing arithmetic libraries, especially from a hardware design perspective. Finally, some directions for further work are outlined. Mikael Henriksson, Theodor Lindberg, Oscar Gustafsson |
ARITH | 3 |
| 2023 | Enhancing Compiler-Driven HDL Design with Automatic Waveform AnalysisabstractThe time-to-market of a new product is one of its most crucial factors for success, therefore, reducing this time is of utter importance. However, this reduction must not come at the expense of a less thorough development process. This paper presents a compiler-driven approach for automatically analyzing metrics such as transaction delays or bus throughput on simulation waveforms of projects developed in the Spade Hardware Description Language (HDL). By utilizing the Spade compiler's knowledge about design internals, an automatic analysis of the waveforms created during simulation is possible using the Waveform Analysis Language (WAL). Analysis programs can be bundled with Spade projects or libraries, such that they are automatically detected by Spade and can be reused by other projects using simple annotations. We call these bundled WAL programs analysis passes, since they fit into the Spade workflow and provide thorough analysis at no additional cost to the users of these libraries. In a detailed description, we present how new analysis passes can be defined using the example of a data streaming interface. Additionally, we highlight the possibilities of analysis passes in two case studies, including Finite State Machine (FSM) and Wishbone protocol analysis. Frans Skarman, Lucas Klemmer, Oscar Gustafsson, Daniel Große |
FDL | 3 |
| 2022 | Spade: An HDL Inspired by Modern Software LanguagesabstractSpade is a new hardware description language which aims to make hardware description easier and less error prone. It does this by taking lessons from software programming languages, and adding language level support for common hardware constructs, all without compromising the low level control over what hardware gets generated. Frans Skarman, Oscar Gustafsson |
FPL | 2 |
| 2021 | Approximate Floating-Point Operations with Integer Units by Processing in the Logarithmic DomainabstractFloating-point numbers represented using a hidden one can readily be approximately converted to the logarithmic domain using Mitchell's approximation. Once in the logarithmic domain, several arithmetic operations including multiplication, division, and square-root can be easily computed using the integer arithmetic unit. This has earlier been used in fast reciprocal square-root algorithms, sometimes referred to as magic number algorithms. The proposed approximate operations are realized by performing an integer operation using an integer unit on floating-point data and adding an integer constant to obtain the approximate floating-point result. In this work, we derive easy to use equations and constants for multiple floating-point formats and operations. Oscar Gustafsson, Noah Hellman |
ARITH | 1 |
| 2021 | Overlap-Save Commutators for High-Speed Streaming Data FilteringabstractOverlap-save and overlap-add methods enable efficient implementation of FIR filters. In this paper, a compact method for handling the overlap and shuffle of samples for real-time processing using pipelined FFT architectures is presented. It is suitable for cases when the sample rate is equal to or higher than the clock frequency. Cheolyong Bae, Oscar Gustafsson |
ISCAS | 2 |
| 2020 | High-Speed Chromatic Dispersion Compensation Filtering in FPGAs for Coherent Optical CommunicationabstractChromatic dispersion is one of the error sources limiting the transmission capacity in coherent optical communication that can be mitigated with digital signal processing. In this paper, the current status and plans of implementation of chromatic dispersion compensation (CDC) filters on FPGAs are discussed. As these high-speed filters are most efficiently implemented in the frequency-domain, different approaches for high-speed FFT-based architectures are considered and preliminary results of fully parallel FFT implementation by utilizing FPGA hardware features are presented. Cheolyong Bae, Oscar Gustafsson |
FPL | 2 |
| 2020 | Acceleration of Simulation Models Through Automatic Conversion to FPGA HardwareabstractBy running simulation models on FPGAs, their execution speed can be significantly improved, at the cost of increased development effort. This paper describes a project to develop a tool which converts simulation models written in high level languages into fast FPGA hardware. The tool currently converts code written using custom C++ data types into Verilog. A model of a hybrid electric vehicle is used as a case study, and the resulting hardware runs significantly faster than on a general purpose CPU. Frans Skarman, Oscar Gustafsson, Daniel Jung 0002, Mattias Krysander |
FPL | 2 |
| 2020 | Pilot-Hopping Sequence Detection Architecture for Grant-Free Random Access using Massive MIMOabstractIn this work, an implementation of a pilot-hopping sequence detector for massive machine type communication is presented. The architecture is based on solution a non-negative least squares problem. The results show that the architecture supporting 1024 users can perform more than one million detections per second with a power consumption of less than 70 mW when implemented in a 28 nm FD-SOI process. Narges Mohammadi Sarband, Ema Becirovic, Mattias Krysander, Erik G. Larsson, Oscar Gustafsson |
ISCAS | 5 |
| 2019 | Direct digital-to-RF converter employing semi-digital FIR voltage-mode RF DAC
M. Reza Sadeghifar, Håkan Bengtsson, J. Jacob Wikner, Oscar Gustafsson |
Integr. | 4 |
| 2019 | Optimum Circuits for Bit-Dimension PermutationsabstractIn this paper, we present a systematic approach to design hardware circuits for bit-dimension permutations. The proposed approach is based on decomposing any bit-dimension permutation into elementary bit-exchanges. Such decomposition is proven to achieve the theoretical minimum number of delays required for the permutation. This offers optimum solutions for multiple well-known problems in the literature that make use of bit-dimension permutations. This includes the design of permutation circuits for the fast Fourier transform, bit reversal, matrix transposition, stride permutations, and Viterbi decoders. Mario Garrido, Jesús Grajal, Oscar Gustafsson |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Karatsuba with Rectangular Multipliers for FPGAsabstractThis work presents an extension of Karatsuba's method to efficiently use rectangular multipliers as a base for larger multipliers. The rectangular multipliers that motivate this work are the embedded 18 × 25-bit signed multipliers found in the DSP blocks of recent Xilinx FPGAs: The traditional Karatsuba approach must under-use them as square 18 × 18 ones. This work shows that rectangular multipliers can be efficiently exploited in a modified Karatsuba method if their input word sizes have a large greatest common divider. In the Xilinx FPG A case, this can be obtained by using the embedded multipliers as 16 × 24 unsigned and as 17 × 25 signed ones. The obtained architectures are implemented with due detail to architectural features such as the pre-adders and post-adders available in Xilinx DSP blocks. They are synthesized and compared with traditional Karatsuba, but also with (non-Karatsuba) state-of-the-art tiling techniques that make use of the full rectangular multipliers. The proposed technique improves resource consumption and performance for multipliers of numbers larger than 64 bits. Martin Kumm, Oscar Gustafsson, Florent de Dinechin, Johannes Kappauf, Peter Zipf |
ARITH | 2 |
| 2017 | On Lifting-Based Fixed-Point Complex Multiplications and RotationsabstractLifting-based complex multiplications and rotations are integer invertible, i.e., an integer input value is mapped to the same integer output value when rotating forward and backward. This is an important aspect for lossless transform based source coding, but since the structure only require three real-valued multiplications and three real-valued additions it is also a potentially attractive way to perform complex multiplications when the coefficient has unity magnitude. In this work, we consider two aspects of these structures. First, we show that both the magnitude and angular error is dependent on the angle of input value and derive both exact and approximated expressions for these. Second, we discuss how to design such structures without the typical separation into three subsequent matrix multiplications. It is shown that the proposed design method allows many more values which are integer invertible, but can not be separated into three subsequent matrix multiplications with fixed-point values. The results show good correspondence between the error approximations and the actual error as well as a significantly increased design space. Oscar Gustafsson |
ARITH | 1 |
| 2017 | Approximate Neumann Series or Exact Matrix Inversion for Massive MIMO?abstractApproximate matrix inversion based on Neumann series has seen a recent increased interest motivated by massive MIMO systems. There, the matrices are in many cases diagonally dominant, and, hence, a reasonable approximation can be obtained within a few iterations of a Neumann series. In this work, we clarify that the complexity of exact methods are about the same as when three terms are used for the Neumann series, so in this case, the complexity is not lower as often claimed. The second common argument for Neumann series approximation, higher parallelism, is indeed correct. However, in most current practical use cases, such a high degree of parallelism is not required to obtain a low latency realization. Hence, we conclude that a careful evaluation, based on accuracy and latency requirements must be performed and that exact matrix inversion is in fact viable in many more cases than the current literature claims. Oscar Gustafsson, Erik Bertilsson, Johannes Klasson, Carl Ingemarsson |
ARITH | 1 |
| 2017 | Efficient FPGA Mapping of Pipeline SDF FFT CoresabstractIn this paper, an efficient mapping of the pipeline single-path delay feedback (SDF) fast Fourier transform (FFT) architecture to field-programmable gate arrays (FPGAs) is proposed. By considering the architectural features of the target FPGA, significantly better implementation results are obtained. This is illustrated by mapping an R22SDF 1024-point FFT core toward both Xilinx Virtex-4 and Virtex-6 devices. The optimized FPGA mapping is explored in detail. Algorithmic transformations that allow a better mapping are proposed, resulting in implementation achievements that by far outperforms earlier published work. For Virtex-4, the results show a 350% increase in throughput per slice and 25% reduction in block RAM (BRAM) use, with the same amount of DSP48 resources, compared with the best earlier published result. The resulting Virtex-6 design sees even larger increases in throughput per slice compared with Xilinx FFT IP core, using half as many DSP48E1 blocks and less BRAM resources. The results clearly show that the FPGA mapping is crucial, not only the architecture and algorithm choices. Carl Ingemarsson, Petter Källström, Fahad Qureshi, Oscar Gustafsson |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Fast and area efficient adder for wide data in recent Xilinx FPGAsabstractMost modern FPGAs have very optimised carry logic for efficient implementations of ripple carry adders (RCA). Some FPGAs also have a six input look up table (LUT) per cell, whereof two inputs are used during normal addition. In this paper we present an architecture that compresses the carry chain length to N/2 in recent Xilinx FPGA, by utilising the LUTs better. This carry compression was implemented by letting some cells calculate the carry chain in two bits per cell, while some others calculate the summary output bits. In total the proposed design uses no more hardware than the normal adder. The result shows that the proposed adder is faster than a normal adder for word length larger than 64 bits in Virtex-6 FPGAs. Petter Källström, Oscar Gustafsson |
FPL | 2 |
| 2016 | Multiplierless Unity-Gain SDF FFTsabstractIn this brief, we propose a novel approach to implement multiplierless unity-gain single-delay feedback fast Fourier transforms (FFTs). Previous methods achieve unity-gain FFTs by using either complex multipliers or nonunity-gain rotators with additional scaling compensation. Conversely, this brief proposes unity-gain FFTs without compensation circuits, even when using nonunity-gain rotators. This is achieved by a joint design of rotators, so that the entire FFT is scaled by a power of two, which is then shifted to unity. This reduces the amount of hardware resources of the FFT architecture, while having high accuracy in the calculations. The proposed approach can be applied to any FFT size, and various designs for different FFT sizes are presented. Mario Garrido, Rikard Andersson, Fahad Qureshi, Oscar Gustafsson |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Decimation filters for high-speed delta-sigma modulators with passband constraints: General versus CIC-based FIR filtersabstractFor high-speed delta-sigma modulators the decimation filters are typically polyphase FIR filters as the recursive CIC filters can not be implemented because of the iteration period bound. In addition, the high clock frequency and short input word length make multiple constant multiplication techniques less beneficial. Instead a realistic complexity measure in this setting is the number of non-zero digits of the FIR filter tap coefficients. As there is limited control of the passband approximation error for CIC-based filters these must in most cases be compensated to meet a passband specification. In this work we investigate the complexity of decimation filters meeting CIC-like stopband behavior, but with a well defined passband approximation error. It is found that the general approach can in many cases produce filters with much smaller passband approximation error at a similar complexity. Oscar Gustafsson, Håkan Johansson |
ISCAS | 1 |
| 2014 | Linear programming design of semi-digital FIR filter and ΣΔ modulator for VDSL2 transmitterabstractAn oversampled digital-to-analog converter including digital ΣΔ modulator and semi-digital FIR filter can be employed in the transmitter of the VDSL2 technology. To select the optimum set of coefficients for the semi-digital FIR filter, an integer optimization problem is formulated in this work, where the model includes the FIR filter magnitude metrics as well as ΣΔ modulator noise transfer function. The semi-digital FIR filter is optimized with respect to magnitude constraints according to the International Telecommunication Union Power Spectral Density mask for VDSL2 technology and minimizing analog cost as the objective function. Utilizing the semi-digital FIR filter with one bit DACs, high linearity required in high-bandwidth profiles of VDSL2, can be achieved. The resolution of the conventional DACs are limited by the mismatch between DAC unit elements. By utilizing one-bit DACs in semi-digital FIR filter, there will be less degradation caused by mismatch between unit elements. The optimization problem is solved in two conditions; fixed passband gain and variable passband gain. It is shown in this paper that 38% saving in total number of unit elements can be achieved by employing variable passband gain in the optimization problem. M. Reza Sadeghifar, J. Jacob Wikner, Oscar Gustafsson |
ISCAS | 3 |
| 2013 | A reconfigurable FFT architecture for variable-length and multi-streaming OFDM standardsabstractThis paper presents a reconfigurable FFT architecture for variable-length and multi-streaming WiMax wireless standard. The architecture processes 1 stream of 2048-point FFT, up to 2 streams of 1024-point FFT or up to 4 streams of 512-point FFT. The architecture consists of a modified radix-2 single delay feedback (SDF) FFT. The sampling frequency of the system is varied in accordance with the FFT length. The latch-free clock gating technique is used to reduce power consumption. The proposed architecture has been synthesized for the Virtex-6 XCVLX760 FPGA. Experimental results show that the architecture achieves the throughput that is required by the WiMax standard and the design has additional features compared to the previous approaches. The design uses 1% of the total available FPGA resources and maximum clock frequency of 313.67 MHz is achieved. Furthermore, this architecture can be expanded to suit other wireless standards. Padma Prasad Boopal, Mario Garrido, Oscar Gustafsson |
ISCAS | 3 |
| 2013 | Low-complexity general FIR filters based on Winograd's inner product algorithmabstractIn this work an FIR filter architecture requiring only half the number of multiplications compared to a direct realization is proposed. The proposed filter architecture is independent on coefficient selection and is therefore suitable for the realization of FIR filters where the filter impulse response is not symmetric/anti-symmetric. The filter architecture is based on a inner product scheme due to Winograd, which to the best of the author's knowledge has not been applied to FIR filters before. A number of different realizations for sequential and two-parallel versions are derived. Oscar Gustafsson, Andreas Ehliar |
ISCAS | 1 |
| 2013 | Pipelined Radix-2k Feedforward FFT ArchitecturesabstractThe appearance of radix-22was a milestone in the design of pipelined FFT hardware architectures. Later, radix-22was extended to radix-2k. However, radix-2kwas only proposed for single-path delay feedback (SDF) architectures, but not for feedforward ones, also called multi-path delay commutator (MDC). This paper presents the radix-2kfeedforward (MDC) FFT architectures. In feedforward architectures radix-2kcan be used for any number of parallel samples which is a power of two. Furthermore, both decimation in frequency (DIF) and decimation in time (DIT) decompositions can be used. In addition to this, the designs can achieve very high throughputs, which makes them suitable for the most demanding applications. Indeed, the proposed radix-2kfeedforward architectures require fewer hardware resources than parallel feedback ones, also called multi-path delay feedback (MDF), when several samples in parallel must be processed. As a result, the proposed radix-2kfeedforward architectures not only offer an attractive solution for current applications, but also open up a new research line on feedforward structures. Mario Garrido, Jesús Grajal, Miguel A. Sánchez Marcos, Oscar Gustafsson |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Using DSP block pre-adders in pipeline SDF FFT implementations in contemporary FPGAsabstractMany contemporary FPGAs have introduced a pre-adder before the hard multipliers, primarily aimed at linear-phase FIR filters. In this work, structural modifications are proposed with the aim of reducing the LUT resource utilization and, finally, using the pre-adder for implementing single path delay feedback pipeline FFTs. The results show that two thirds of the LUT resources can be saved when the pre-adder has bypass functionality, as in the Xilinx 6 and 7 series, compared to a direct mapping. Carl Ingemarsson, Petter Källström, Oscar Gustafsson |
FPL | 3 |
| 2011 | Implementation of time-multiplexed sparse periodic FIR filters for FRM on FPGAsabstractFrequency-response masking (FRM) is a set of techniques for lowering the computational complexity of narrow transition band FIR filters. These FRM use a combination of sparse periodic filters and non-sparse filters. In this work we consider the implementation of these filters in a time-multiplexed manner on FPGAs. It is shown that the proposed architectures produce lower complexity realizations compared to the vendor provided IP blocks, which do not take the sparseness into consideration. The designs are implemented on a Virtex-6 device utilizing the built-in DSP blocks. Syed Asad Alam, Oscar Gustafsson |
ISCAS | 2 |
| 2011 | Minimum adder depth multiple constant multiplication algorithm for low power FIR filtersabstractIn this work we propose a graph based minimum adder depth algorithm for the multiple constant multiplication (MCM) problem. Hence, all multiplier coefficients are here guaranteed to be realized at the theoretically lowest depth possible. The motivation for low adder depth is that this has been shown to be a main factor for the power consumption. An FIR filter is implemented using different MCM algorithms, and the proposed algorithm result in 25% lower power in the MCM part compared to algorithms focused on minimizing the number of adders. Kenny Johansson, Oscar Gustafsson, Linda DeBrunner, Lars Wanhammar |
ISCAS | 2 |
| 2010 | Redundancy reduction for high-speed fir filter architectures based on carry-save adder treesabstractIn this work we consider high-speed FIR filter architectures implemented using, possibly pipelined, carry-save adder trees for accumulating the partial products. In particular we focus on the mapping between partial products and full adders and propose a technique to reduce the number of carry-save adders based on the inherent redundancy of the partial products. The redundancy reduction is performed on the bit-level to also work for short wordlength data such as those obtained from sigma-delta modulators. Anton Blad, Oscar Gustafsson |
ISCAS | 2 |
| 2010 | Twiddle factor memory switching activity analysis of radix-22 and equivalent FFT algorithmsabstractIn this paper, we propose equivalent radix-22algorithms and evaluate them based on twiddle factor switching activity for a single delay feedback pipelined FFT architecture. These equivalent pipeline FFT algorithms have the same number of complex multipliers with the same resolution as the radix-22. It is shown that the twiddle factor switching activity of the equivalent algorithms is reduced with up to 40% for some of the equivalent algorithms derived for N = 256. Fahad Qureshi, Oscar Gustafsson |
ISCAS | 2 |
| 2010 | Addition Aware Quantization for Low Complexity and High Precision Constant MultiplicationabstractMultiplication by constants can be efficiently realized using shifts, additions, and subtractions. In this work we consider how to select a fixed-point value for a real valued, rational, or floating-point coefficient to obtain a low-complexity realization. It is shown that the process, denoted addition aware quantization, often can determine coefficients that has as low complexity as the rounded value, but with a smaller approximation error by searching among coefficients with a longer wordlength. Oscar Gustafsson, Fahad Qureshi |
IEEE Signal Process. Lett. | 1 |
| 2009 | Estimation of the Switching Activity in Shift-and-add based ComputationsabstractIn this work, we propose a switching activity model for constant multipliers. The model can also be used for other architectures that are composed by full adders. Hence, the proposed model is suitable to be used in power consumption aware design algorithms. An important category is algorithms for the multiple-constant multiplication (MCM) problem. The model is shown to agree well with simulations, especially for carry-save arithmetic. Kenny Johansson, Oscar Gustafsson, Linda DeBrunner |
ISCAS | 2 |
| 2009 | Scaling of Fractional Delay Filters based on the Farrow StructureabstractIn this work we consider scaling of fractional delay filters using the Farrow structure. Based on the observation that the subfilters approximate the Taylor expansion of a differentiator, we derive estimates of the L2-norm scaling values at the outputs of each subfilter as well as at the inputs of each delay multiplier. The scaling values can then be used to derive suitable wordlengths in a fixed-point implementation. Tahir Abbas Khan, Oscar Gustafsson, Håkan Johansson |
ISCAS | 2 |
| 2009 | Low-complexity Reconfigurable Complex Constant Multiplication for FFTsabstractIn this work we consider structures for simultaneous multiplication by a small set of two pairwise coefficients where the coefficients are the real and imaginary part of a limited number of points uniformly spread on the unit circle. Hence, each such multiplier forms half of a complex multiplier suitable for twiddle factor multiplication in FFT architectures. Based on trigonometric identities we propose a multiplier for a unit circle resolution of 32 points. Also, we revisit an earlier proposed multiplier for 16 points and show that the complexity can be reduced by using minimum adder constant multipliers compared with the earlier proposed CSD-based multipliers. Fahad Qureshi, Oscar Gustafsson |
ISCAS | 2 |
| 2008 | Bit-level optimized FIR filter architectures for high-speed decimation applicationsabstractAnalog-to-digital converters based on sigma-delta modulation have shown promising performance, with steadily increasing bandwidth. However, associated with the increasing bandwidth is an increasing output sample rate, which becomes costly to decimate in the digital domain. Commonly, cascaded integrator comb structures have been used for the first decimation stage, but polyphase decomposed FIR filter architectures have been shown to be more power efficient. In this paper, a bit-level optimization algorithm is introduced, and applied to the direct form and transposed form FIR filter architectures. Mainly, two conclusions can be drawn. The transposed architecture has significantly lower complexity in most circumstances, and the inability to implement an efficient adder prohibits the symmetry of the filter coefficients to be used efficiently for the direct form architecture. Anton Blad, Oscar Gustafsson |
ISCAS | 2 |
| 2008 | Switching activity estimation for shift-and-add based constant multipliersabstractIn this work we propose a switching activity model for single adder multipliers. This correspond to the case where a signal is added to a shifted version of itself, which is a common part in multiple constant multiplication (MCM). Hence, the proposed model is suitable to be used in power consumption aware MCM algorithms. The model is shown to agree well with simulations, and for the studied test cases a maximum error of 0.26% is obtained. Kenny Johansson, Oscar Gustafsson, Lars Wanhammar |
ISCAS | 2 |
| 2008 | Power optimization of weighted bit-product summation tree for elementary function generatorabstractIn this paper we propose a method for lowering the power consumption in our previously proposed method for approximating elementary functions. By rearranging the interconnect ordering in the summation tree we show that it is possible to lower the power consumption in the range of 5.4% to 25.6% compared to a random ordering. The reduction tree is progressively designed and the interconnect ordering is decided based on the transition activities of the partial products. The reduction in power consumption comes with no overhead in performance or area compared to the random ordering. Saeeid Tahmasbi Oskuii, Kenny Johansson, Oscar Gustafsson, Per Gunnar Kjeldsberg |
ISCAS | 3 |
| 2007 | Transition-activity aware design of reduction-stages for parallel multipliersabstractWe propose an interconnect reorganization algorithm for reduction stages in parallel multipliers. It aims at minimizing power consumption for given static probabilities at the primary inputs. In typical signal processing applications the transition probability varies between the most and least significant bits. The same is the case for individual signals within the multiplier. Our interconnect reorganization exploits this to reduce the overall switching activity, thus reducing the multiplier's power consumption. We have developed a CAD tool that reorganizes the connections within the multiplier architecture in an optimized way. Since the applied heuristic requires power estimation, we have also developed a very fast estimator fine tuned for parallel multipliers. The CAD tool automatically generates gate-level VHDL code for the optimizedmultipliers. This code and code for unoptimized multipliers have been compared using state of the art power estimation tools. The reduction in power consumption ranges from 7% up to 23% and can be achieved without any noticeable overhead in performance and area. Saeeid Tahmasbi Oskuii, Per Gunnar Kjeldsberg, Oscar Gustafsson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | A Difference Based Adder Graph Heuristic for Multiple Constant Multiplication ProblemsabstractMultiple constant multiplication (MCM), i.e., realizing a number of constant multiplications using a minimum number of adders and subtracters, has been an active research area for the last decade. An adder graph type algorithm for solving the MCM problem is introduced with a novel heuristic inspired by difference methods. It is shown that the results is as good or better as previous state of the art under most conditions. Furthermore, the proposed algorithm does not rely on look-up tables. Oscar Gustafsson |
ISCAS | 1 |
| 2007 | Complexity Comparison of Linear-Phase Mth-Band and General FIR FiltersabstractLinear-phaseMth-band finite-length impulse response (FIR) filters are characterized by everyMth impulse response coefficient being equal to zero, except for the middle tap. This leads to a realization with few multiplications. However, other properties include that the passband ripple is bounded by the stopband ripple and that the transition band is symmetric around π/Mrad. Depending on the filter specifications this may sometimes lead to an overly constrained filter design. In this work we investigate the trade-off betweenMth-band and general linear-phase FIR filters. The results are bounds on the filter specification for when the linear-phaseMth-band or general FIR filter is advantageous to use. Oscar Gustafsson, Håkan Johansson |
ISCAS | 1 |
| 2007 | Complexity Reduction of Constant Matrix Computations over the Binary Field
Oscar Gustafsson, Mikael Olofsson |
WAIFI | 1 |
| 2006 | Bidirectional conversion to minimum signed-digit representationabstractIn this work an approach to converting a number in two's complement representation to a minimum signed-digit representation is proposed. The novelty in this work is that this conversion is done from left-to-right and right-to-left concurrently. Hence, the execution time is significantly decreased, while the area overhead is small. Erik Backenius, Erik Säll, Oscar Gustafsson |
ISCAS | 3 |
| 2006 | Approximation of elementary functions using a weighted sum of bit-productsabstractIn this work a novel approach for approximating elementary functions is presented. By rewriting the function as a sum of weighted bit-products an efficient implementation is obtained. For most functions a majority of the bit-products can be neglected and still obtain good accuracy. The method is suitable for high-speed implementation of fixed-point functions Kenny Johansson, Oscar Gustafsson, Lars Wanhammar |
ISCAS | 2 |
| 2000 | Design and efficient implementation of high-speed narrow-band recursive digital filters using single filter frequency masking techniquesabstractIn this paper a novel structure for frequency masking based on a single filter is introduced. By using the proposed frequency masking approach the maximal sample frequency for narrow-band recursive digital filters is increased. This increase in sample frequency can be utilized either for faster filters or for lower power consumption by trading excess speed for low power. As the subfilters are identical (except for the number of delays), the filter can be mapped onto a single hardware structure which leads to area-efficient implementation. Oscar Gustafsson, Håkan Johansson, Lars Wanhammar |
ISCAS | 1 |