Renato J. Cintra

dblp:51/7601 · DBLP profile ↗
← Back
45ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-4579-6757ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 5 since 2021Systems, architecture and hardware · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Low-Complexity Approximations of the WECS Method for SAR Change Detection
abstract
This letter introduces low-complexity approximations for the Wavelet Energy Correlation Screening (WECS) method, which aims at change detection in multitemporal SAR images. The WECS method relies on the non-decimated discrete wavelet transform (ND-DWT) to compute approximation coefficients employed in a feature screening process based on the Pearson correlation. Although effective, WECS presents a high computational cost due to its repeated wavelet filtering stage. To overcome this drawback, we propose two approximations for the wavelet filter coefficients, obtained by truncating their canonical signed digit (CSD) representation, which significantly reduces the number of arithmetic operations. Numerical experiments using both simulated and real-world datasets demonstrate that the proposed methods not only maintain the performance of the original WECS but also achieve computational gains, even outperforming it in certain scenarios.
Luan Portella, Renato J. Cintra, Aluísio Pinheiro
IEEE Signal Process. Lett.2
2025 Low-Complexity Combined Approximate DFT and Adaptive Beamformer for Extremely Large Arrays
abstract
Discrete Fourier Transform (DFT) based beam-formers provide N beams with a uniform linear array having N antennas. Despite low-complexity, DFT beamformers cannot generate deep nulls at arbitrary directions of arrivals (DOAs) in the beam pattern to attenuate interferences. On the other hand, data-dependent adaptive beamformers can achieve deep nulls at arbitrary DOAs, however, at higher complexity due to an inversion of a matrix in deriving optimum weights. In this paper, we propose a combined approximate DFT (ADFT) and adaptive beamformer for a uniform linear array consisting of P subarrays, each consisting of M antennas. Here, we first employ P ADFT beamformers to process signals received by subarrays, and the P outputs of these subarrays are then processed by an adaptive beamformer. Simulation results confirm that the proposed beamformer provides the benefits of both DFT and adaptive beamformers. Furthermore, we present preliminary results of an experimental antenna array operating at 5.75 GHz, where a Howells-Applebaum beamformer is employed as the adaptive beamformer.
Arjuna Madanayake, Umesha Kumarasiri, S. Sivasankar, Chamira U. S. Edussooriya, Renato J. Cintra, Chamith Wijenayake
ISCAS5
2025 Analog-Digital Approximate DFT with Spatial ∆-Σ LNA Multi-beam RF Apertures
abstract
Multifunctional and waveform agnostic antenna apertures having multiple simultaneous RF beams are necessary for emerging electromagnetic situational awareness applications. A multibeam aperture operating across the entire 1-6 GHz band is crucial for both military and commercial wireless applications. Spectrum perception refers to the application of artificial intelligence (AI) and machine learning (ML) to wireless applications to extract intelligence on spectral activity as a function of both direction and waveform parameters (frequency, modulation, and waveform shape). Efficient and accurate RF sensing is a necessary precursor to higher level AI/ML spectral perception algorithms for detecting particular waveforms, behaviors, or signatures. This paper explores multibeam beamforming for RF sensing as a joint analog-digital hybrid approximate computing problem that can be efficiently implemented using multiple chiplets. The first chiplet includes both a multiport LNA with Δ-Σ spatial noise shaping to improve resilience to high power jammers, and an approximate DFT (ADFT) based analog multi-beamformer with reduced circuit complexity. The second chiplet includes ADCs and ADFT-based digital beamformers with reduced computational complexity compared to conventional DFT-based designs. Initial results on the design of both chiplets are presented.
Arjuna Madanayake, H. Pilippange, Keththura Lawrance, A. Uddin, S. Mandal, J. Di, M. Tennant, C. Workman, Renato J. Cintra
ISCAS9
2025 Area-Efficient FPGA Architectures for Multidimensional DCT using Approximate Transforms and Computing
abstract
The combined use of algorithms for approximate discrete cosine transform (DCT) with approximate computing is examined toward area-efficient field programmable gate array (FPGA) hardware architectures. The performance of previously reported, multiplierless approximate DCT transforms are explored in the presence of non-exact hardware adders which provide better utilisation of FPGA resources in terms of look-up-tables (LUTs). In the context of 2D image compression, provided examples are demonstrating a 20% LUT reduction in FPGA utilization compared to accurate adders, with a corresponding image quality degradation of 0.6 dB in terms of PSNR. The potential benefit of such area-efficient FPGA hardware designs is explored on 3D (video compression), 4D (static light field compression), and 5D (light field video compression) cases where significant FPGA resource savings are obtained at the cost of marginal degradation of output quality.
Pathmapirian Nanthakumar, Chamith Wijenayake, Chamira U. S. Edussooriya, Arjuna Madanayake, Renato J. Cintra
ISCAS5
2025 Real-Time 5.7-5.8 GHz 32-Beam Approximate Discrete Fourier Transform Spectrum Sensor for RF Perception on Xilinx Sx475T
abstract
The radio spectrum in the sub-6 GHz (FR1) band is crowded and contested, and is sought after by commercial, scientific and defense users. Situational awareness through spectrum sensing, and AI/ML-enabled perception that recognizes behaviors, patterns, modulations, devices and waveforms is a crucial need for emerging autonomous/cognitive radio systems. This work describes measurable progress in the use of extremely low complexity approximate DFT algorithms as multi-beam beamformers in the digital domain for multibeam spatial RF beamforming. The paper begins with a longterm vision for intelligent spectrum awareness across wide bands and multi-directions with multi-chiplet system in package hardware acceleration of both beamforming, Fourier and AI/ML algorithms, followed by a focus account of specific progress with digital architectures and real-time prototype implementations across the 5.7–5.8 GHz band for 32 RF beams. A real-time temporal frequency resolution of 100 kHz across 100 MHz of baseband bandwidth is achieved, across 32 simultaneous fully-digital RF-beams, using a Xilinx Sx475 FPGA implementation. Details of multiplierless approximate DFT beamformers, automated modulation recognition algorithms using AI/ML, analog channelization, spectrum sensing and perception architectures are also discussed. Over-the-air experiments using the RadioML.2018.a dataset confirmed both single source accuracy (better than 97%) and impact of multi-beams on AI/ML performance for multiple strong RFI sources.
Arjuna Madanayake, Umesha Kumarasiri, Sivakumar Sivasankar, Keththura Lawrance, Buddhipriya Gayanath, Hiruni Silva, Soumyajit Mandal, Renato J. Cintra
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 Fast data-independent KLT approximations based on integer functions
Anabeth P. Radünz, Diego F. G. Coelho, Fábio M. Bayer, Renato J. Cintra, Arjuna Madanayake
Multim. Tools Appl.4
2022 Improved Point Estimation for the Rayleigh Regression Model
abstract
The Rayleigh regression model was recently proposed for modeling amplitude values of synthetic aperture radar (SAR) image pixels. However, inferences from such model are based on the maximum-likelihood estimators, which can be biased for small-signal lengths. The Rayleigh regression model for SAR images often considers small pixel windows, which may lead to inaccurate results. In this letter, we introduce bias-adjusted estimators tailored for the Rayleigh regression model based on: 1) Cox and Snell’s method; 2) Firth’s scheme; and 3) the parametric bootstrap method. We present numerical experiments considering synthetic and actual SAR data sets. The bias-adjusted estimators yield nearly unbiased estimates and accurate modeling results.
Bruna G. Palm, Fábio M. Bayer, Renato J. Cintra
IEEE Geosci. Remote. Sens. Lett.3
2022 Data-independent low-complexity KLT approximations for image and video coding
Anabeth P. Radünz, Thiago L. T. da Silveira, Fábio M. Bayer, Renato J. Cintra
Signal Process. Image Commun.4
2022 Radix-$N$ Algorithm for Computing $N^{2^{n}}$-Point DFT Approximations
abstract
The ever increasing technological demand for the DFT computation poses several challenges both to theory and hardware realization. The design of usual fast Fourier transform (FFT) algorithms seems to have reached a stage of diminishing returns in terms of performance. Alternatively, approximate transform methods have been demonstrated to provide substantial gains in terms of energy-efficiency and performance by tolerating small inaccuracies in the results. In this paper, we present a transform scaling method variant of the Cooley-Tukey algorithm to obtain DFT approximations of large blocksize. The proposed method scales up a givenN-point transformation to anN2-point transformation. Such scaling can be successively applied leading to$\mathop {{N}^2}\nolimits^n $-point transformations. We have fully presented the 324-point DFT approximation which stems from a multiplierless 32-point DFT approximation. The proposed approximation is equipped with a fast algorithm; we also supply the arithmetic complexity assessment and an error analysis.
Luan Portella, Diego F. G. Coelho, Fábio M. Bayer, Arjuna Madanayake, Renato J. Cintra
IEEE Signal Process. Lett.5
2022 A Class of Low-Complexity DCT-Like Transforms for Image and Video Coding
abstract
The discrete cosine transform (DCT) is a relevant tool in signal processing applications, mainly known for its good decorrelation properties. Current image and video coding standards—such as JPEG and HEVC—adopt the DCT as a fundamental building block for compression. Recent works have introduced low-complexity approximations for the DCT, which become paramount in applications demanding real-time computation and low-power consumption. The design of DCT approximations involves a trade-off between computational complexity and performance. This paper introduces a new multiparametric transform class encompassing the round-off DCT (RDCT) and the modified RDCT (MRDCT), two relevant multiplierless 8-point approximate DCTs. The associated fast algorithm is provided. Four novel orthogonal low-complexity 8-point DCT approximations are obtained by solving a multicriteria optimization problem. The optimal 8-point transforms are scaled to lengths 16 and 32 while keeping the arithmetic complexity low. The proposed methods are assessed by proximity and coding measures with respect to the exact DCT. Image and video coding experiments and hardware realization are performed. The novel transforms perform close to or outperform the current state-of-the-art DCT approximations.
Thiago L. T. da Silveira, Diego Ramos Canterle, Diego F. G. Coelho, Vítor de A. Coutinho, Fábio M. Bayer, Renato J. Cintra
IEEE Trans. Circuits Syst. Video Technol.6
2022 Robust Rayleigh Regression Method for SAR Image Processing in Presence of Outliers
abstract
The presence of outliers (anomalous values) in synthetic aperture radar (SAR) data and the misspecification in statistical image models may result in inaccurate inferences. To avoid such issues, the Rayleigh regression model based on a robust estimation process is proposed as a more realistic approach to model this type of data. This article aims at obtaining Rayleigh regression model parameter estimators robust to the presence of outliers. The proposed approach considered the weighted maximum likelihood method and was submitted to numerical experiments using simulated and measured SAR images. Monte Carlo simulations were employed for the numerical assessment of the proposed robust estimator performance in finite signal lengths, their sensitivity to outliers, and the breakdown point. For instance, the nonrobust estimators show a relative bias value 65-fold larger than the results provided by the robust approach in corrupted signals. In terms of sensitivity analysis and break down point, the robust scheme resulted in a reduction of about 96% and 10%, respectively, in the mean absolute value of both measures, in compassion to the nonrobust estimators. Moreover, two SAR datasets were used to compare the ground type and anomaly detection results of the proposed robust scheme with competing methods in the literature.
Bruna G. Palm, Fábio M. Bayer, Renato B. Machado, Mats I. Pettersson, Viet Thuy Vu, Renato J. Cintra
IEEE Trans. Geosci. Remote. Sens.6
2020 Low-Complexity Real-Time Light Field Compression using 4-D Approximate DCT
abstract
A low-complexity codec and a hardware architecture are proposed for achieving real-time compression of four-dimensional (4-D) light field (LF) signals captured from camera/lenslet arrays. The proposed system employs the 4-D extension of the two-dimensional (2-D) 8×8 approximate discrete cosine transform (ADCT) that has recently appeared in the literature. Motivated by the partial separability of the multidimensional spectrum of LFs, the proposed 4-D ADCT is obtained by cascading 2-D inter-view and 2-D intra-view transform stages. Software simulations are provided to confirm the performance of the 4-D ADCT based compression and comparisons are made with respect to 2-D inter-view only and 2-D intra-view only ADCT-based compression. Proposed digital architectures are validated using stepped hardware co-simulation on a Xilinx Virtex-7 VC-707 FPGA platform verifying 597 MHz maximum possible clock frequency, implying an ideal throughput of 18×103LFs/sec for performing 4-D ADCT on (8× 8×432×624×3) size LFs. When 10% of the ADCT coefficients per each (8×8×8×8) hypercube are retained sub aperture images show 38 dB average PSNR and 0.95 average SSIM.
Namalka Liyanage, Chamith Wijenayake, Chamira U. S. Edussooriya, Arjuna Madanayake, Renato J. Cintra, Eliathamby Ambikairajah
ISCAS5
2020 A Multiparametric Class of Low-complexity Transforms for Image and Video Coding
Diego Ramos Canterle, Thiago L. T. da Silveira, Fábio M. Bayer, Renato J. Cintra
Signal Process.4
2019 Rayleigh Regression Model for Ground Type Detection in SAR Imagery
abstract
This letter proposes a regression model for nonnegative signals. The proposed regression estimates the mean of Rayleigh distributed signals by a structure which includes a set of regressors and a link function. For the proposed model, we present: 1) parameter estimation; 2) large data record results; and 3) a detection technique. In this letter, we present closed-form expressions for the score vector and Fisher information matrix. The proposed model is submitted to extensive Monte Carlo simulations and to the measured data. The Monte Carlo simulations are used to evaluate the performance of maximum likelihood estimators. Also, an application is performed comparing the detection results of the proposed model with Gaussian-, Gamma-, and Weibull-based regression models in synthetic aperture radar (SAR) images.
Bruna G. Palm, Fábio M. Bayer, Renato J. Cintra, Mats I. Pettersson, Renato B. Machado
IEEE Geosci. Remote. Sens. Lett.3
2019 An iterative wavelet threshold for signal denoising
Fábio M. Bayer, Alice J. Kozakevicius, Renato J. Cintra
Signal Process.3
2019 Detecting Changes in Fully Polarimetric SAR Imagery With Statistical Information Theory
abstract
Images obtained from coherent illumination processes are contaminated with speckle. A prominent example of such imagery systems is the polarimetric synthetic aperture radar (PolSAR). For such a remote sensing tool, the speckle interference pattern appears in the form of a positive-definite Hermitian matrix, which requires specialized models and makes change detection a hard task. The scaled complex Wishart distribution is a widely used model for PolSAR images. Such a distribution is defined by two parameters: the number of looks and the complex covariance matrix. The last parameter contains all the necessary information to characterize the backscattered data, and thus, identifying changes in a sequence of images can be formulated as a problem of verifying whether the complex covariance matrices differ at two or more takes. This paper proposes a comparison between a classical change detection method based on the likelihood ratio and three statistical methods that depend on information-theoretic measures: the Kullback-Leibler (KL) distance and two entropies. The performance of these four tests was quantified in terms of their sample test powers and sizes using simulated data. The tests are then applied to actual PolSAR data. The results provide evidence that tests based on entropies may outperform those based on the KL distance and likelihood ratio statistics.
Abraao D. C. Nascimento, Alejandro C. Frery, Renato J. Cintra
IEEE Trans. Geosci. Remote. Sens.3
2018 An Offset-Canceling Approximate-DFT Beamforming Architecture for Wireless Transceivers
abstract
We describe a current-mode multi-beam beamforming approach for 5G wireless applications based on a low-complexity approximate-DFT (a-DFT). Dynamic current mirrors are used to cancel errors in the current copying and scaling operations required to realize a-DFT matrices, thus resulting in an accurate and scalable architecture. The circuit design for the case of 8-point a-DFT has been validated with transistor-level simulations in the UMC 0.18 μm CMOS process.
Haixiang Zhao, Soumyajit Mandal, Viduneth Ariyarathna, Arjuna Madanayake, Renato J. Cintra
ISCAS5
2018 Computation of 2D 8×8 DCT Based on the Loeffler Factorization Using Algebraic Integer Encoding
abstract
This paper proposes a computational method for 2D 8×8 DCT based on algebraic integers. The proposed algorithm is based on the Loeffler 1D DCT algorithm, and it is shown to operate with exact computation—i.e., error-free arithmetic—up to the final reconstruction step (FRS). The proposed algebraic integer architecture maintains error-free computations until an entire block of DCT coefficients having size 8×8 is computed, unlike algorithms in the literature which claim to be error-free but in fact introduce arithmetic errors between the column- and row-wise 1D DCT stages in a 2D DCT operation. Fast algorithms are proposed for the final reconstruction step employing two approaches, namely, the expansion factor and dyadic approximation. A digital architecture is also proposed for a particular FRS algorithm, and is implemented on an FPGA platform for on-chip verification. The FPGA implementation operates at 360 MHz, and is capable of a real-time throughput of$3.6\cdot 10^8$2D DCTs of size 8×8 every second, with corresponding pixel rate of$2.3\cdot 10^{10}$pixels per second. The digital architecture is synthesized using 180 nm CMOS standard cells and shows a chip area of 7.41 mm$^2$. The CMOS design is predicted to operate at 893 MHz clock frequency, at a dynamic power consumption 13.22 mW/MHz$\cdot$V$_{sup}^2$.
Diego F. G. Coelho, Sushmabhargavi Nimmalapalli, Vassil S. Dimitrov, Arjuna Madanayake, Renato J. Cintra, Arnaud Tisserand
IEEE Trans. Computers5
2018 Low-Complexity Approximate Convolutional Neural Networks
abstract
In this paper, we present an approach for minimizing the computational complexity of the trained convolutional neural networks (ConvNets). The idea is to approximate all elements of a given ConvNet and replace the original convolutional filters and parameters (pooling and bias coefficients; and activation function) with an efficient approximations capable of extreme reductions in computational complexity. Low-complexity convolution filters are obtained through a binary (zero and one) linear programming scheme based on the Frobenius norm over sets of dyadic rationals. The resulting matrices allow for multiplication-free computations requiring only addition and bit-shifting operations. Such low-complexity structures pave the way for low power, efficient hardware designs. We applied our approach on three use cases of different complexities: 1) a "light" but efficient ConvNet for face detection (with around 1000 parameters); 2) another one for hand-written digit classification (with more than 180 000 parameters); and 3) a significantly larger ConvNet: AlexNet with million matrices. We evaluated the overall performance on the respective tasks for different levels of approximations. In all considered applications, very low-complexity approximations have been derived maintaining an almost equal classification performance.
Renato J. Cintra, Stefan Duffner, Christophe Garcia, André Leite
IEEE Trans. Neural Networks Learn. Syst.1
2017 A Parallel Method for the Computation of Matrix Exponential Based on Truncated Neumann Series
abstract
This paper introduces a new method for computing matrix exponential based on truncated Neumann series. The efficiency of the method is based on smart factorizations for evaluation of several Neumann series that can be done in parallel and divided across different processors with low communication overhead. A physical realization on FPGA is provided for proof-of-concept. The method is verified to be advantageous over the usual Horner's rule approach for polynomial evaluation. The hardware verification shows a reduction of 62% in time required for processing for series approximations with 9 terms. Software verification demonstrates a 30% reduction in time compared to Horner's rule and the trade-offs between using a higher precision approach is illustrated.
Vassil S. Dimitrov, Viduneth Ariyarathna, Diego F. G. Coelho, Logan Rakai, Arjuna Madanayake, Renato J. Cintra
ARITH6
2017 DCT approximations based on Chen's factorization
C. J. Tablada, Thiago L. T. da Silveira, Renato J. Cintra, Fábio M. Bayer
Signal Process. Image Commun.3
2017 DFT Computation Using Gauss-Eisenstein Basis: FFT Algorithms and VLSI Architectures
abstract
A joint numerical representation based on both Gaussian and Eisenstein integers is proposed. This Gauss-Eisenstein representation maps complex numbers into four-tuples of integers with arbitrarily high precision. The representation furnishes the computation of the 3-, 6-, and 12-point discrete Fourier transform (DFT) at any desired accuracy. The associated fast algorithms based on the Gauss-Eisenstein integers are error-free up to the final reconstruction step, which can be realized in hardware as a multiplierless implementation. The introduced methods are compared with competing algorithms in terms of arithmetic complexity. We propose three FRS architectures based on the following methods: Dempster-McLeod representation, expansion factor, and addition aware quantization. The Gauss-Eisenstein 12-point DFT is physically realized on a Xilinx Virtex 6 FPGA device with maximum clock frequency of 302 MHz for the expansion factor FRS with real-time throughput of 3:62 × 109 coefficients/s. The FPGA verified digital designs were synthesized, mapped, placed and finally routed for 0:18mm CMOS technology assuming a 1.8 V DC supply employing Austria Micro Systems (AMS) standard-cell library (hitkit version 4.11). The routed ASIC is predicted to operate at a maximum frequency of 505 MHz for the expansion factor FRS with potential real-time throughput of 6:06 × 109coefficients/s.
Diego F. G. Coelho, Renato J. Cintra, Nilanka T. Rajapaksha, Gihan J. Mendis, Arjuna Madanayake, Vassil S. Dimitrov
IEEE Trans. Computers2
2017 Low-Complexity Image and Video Coding Based on an Approximate Discrete Tchebichef Transform
abstract
The usage of linear transformations has great relevance for data decorrelation applications, like image and video compression. In that sense, the discrete Tchebichef transform (DTT) possesses useful coding and decorrelation properties. The DTT transform kernel does not depend on the input data and fast algorithms can be developed to real-time applications. However, the DTT fast algorithm presented in literature possess high computational complexity. In this paper, we introduce a new low-complexity approximation for the DTT. The fast algorithm of the proposed transform is multiplication free and requires a reduced number of additions and bit-shifting operations. Image and video compression simulations in popular standards show good performance of the proposed transform. Regarding hardware resource consumption for FPGA shows a 43.1% reduction in configurable logic blocks and ASIC place and route realization shows a 57.7% reduction in the area-time figure compared with the 2D version of the exact DTT.
Paulo A. M. Oliveira, Renato J. Cintra, Fábio M. Bayer, Sunera Kulasekera, Arjuna Madanayake
IEEE Trans. Circuits Syst. Video Technol.2
2017 Low-Complexity Multidimensional DCT Approximations for High-Order Tensor Data Decorrelation
abstract
In this paper, we introduce low-complexity multidimensional discrete cosine transform (DCT) approximations. 3D DCT approximations are formalized in terms of high-order tensor theory. The formulation is extended to higher dimensions with arbitrary lengths. Several multiplierless 8×8 ×8 approximate methods are proposed and the computational complexity is discussed for the general multidimensional case. The proposed methods complexity cost was assessed, presenting considerably lower arithmetic operations when compared with the exact 3D DCT. The proposed approximations were embedded into 3D DCT-based video coding scheme and a modified quantization step was introduced. The simulation results showed that the approximate 3D DCT coding methods offer almost identical output visual quality when compared with exact 3D DCT scheme. The proposed 3D approximations were also employed as a tool for visual tracking. The approximate 3D DCT-based proposed system performs similarly to the original exact 3D DCT-based method. In general, the suggested methods showed competitive performance at a considerably lower computational cost.
Vítor de A. Coutinho, Renato J. Cintra, Fábio M. Bayer
IEEE Trans. Image Process.2
2016 Error-free computation of 8-point discrete cosine transform based on the Loeffler factorisation and algebraic integers
abstract
An 8‐point discrete cosine transform (DCT) fast algorithm based on the Loeffler DCT factorisation and algebraic integer (AI) representation is proposed. The proposed algorithm is an error‐free implementation of the Loeffler algorithm and it is capable of computing the 8‐point DCT multiplierlessly. Decoding architectures are also proposed for mapping AI encoded quantities back to usual fixed point arithmetic using canonical signed digit representation and the expansion factor method. The proposed algorithm is mapped into systolic‐array digital architectures and physically realised as digital prototype circuits using field‐programmable gate array technology on a Reconfigurable Open Architecture Computing Hardware board and mapped to 0.18 μm complementary metal–oxide–semiconductor technology using AMS Encounter Digital Implementation libraries at 1.8 V supply.
Diego F. G. Coelho, Renato J. Cintra, Sunera Kulasekera, Arjuna Madanayake, Vassil S. Dimitrov
IET Signal Process.2
2015 Fast computation of residual complexity image similarity metric using low-complexity transforms
abstract
The authors apply two approaches to reduce the computation time of the residual complexity similarity metric employed in image registration applications aimed at hardware‐based implementations with low‐complexity transforms. First, the similarity metric is computed in image sub‐blocks, which are subsequently combined into a global metric value. Second, the discrete cosine transform (DCT) needed in the computation of the similarity measure is replaced with multiplier‐free low‐complexity approximate transforms. The authors propose a new low‐complexity transform requiring only 18 additions in an 8 × 8 block and compare it to: the round DCT, the signed DCT, the Hadamard transform and the Walsh‐Hadamard transform. Detailed computational complexity analysis reveals that block‐wise processing alone reduces computational cost by a factor of 8‐9 for original DCT composed of multiplications and additions, and up to ≃4.90 when the proposed DCT is utilised; being the computation performed with additions only. Results obtained from computer simulated and realistic X‐ray images demonstrate block‐wise processing and approximate transforms result in successful image registration, making residual complexity similarity measure available to hardware‐accelerated fast image registration applications.
Yves Pauchard, Renato J. Cintra, Arjuna Madanayake, Fábio M. Bayer
IET Image Process.2
2015 A class of DCT approximations based on the Feig-Winograd algorithm
C. J. Tablada, Fábio M. Bayer, Renato J. Cintra
Signal Process.3
2015 A Discrete Tchebichef Transform Approximation for Image and Video Coding
abstract
In this letter, we introduce a low-complexity approximation for the discrete Tchebichef transform (DTT). The proposed forward and inverse transforms are multiplication-free and require a reduced number of additions and bit-shifting operations. Numerical compression simulations demonstrate the efficiency of the proposed transform for image and video coding. Furthermore, Xilinx Virtex-6 FPGA based hardware realization shows 44.9% reduction in dynamic power consumption and 64.7% lower area when compared to the literature.
Paulo A. M. Oliveira, Renato J. Cintra, Fábio M. Bayer, Sunera Kulasekera, Arjuna Madanayake
IEEE Signal Process. Lett.2
2015 VLSI Computational Architectures for the Arithmetic Cosine Transform
abstract
The discrete cosine transform (DCT) is a widely-used and important signal processing tool employed in a plethora of applications. Typical fast algorithms for nearly-exact computation of DCT require floating point arithmetic, are multiplier intensive, and accumulate round-off errors. Recently proposed fast algorithm arithmetic cosine transform (ACT) calculates the DCT exactly using only additions and integer constant multiplications, with very low area complexity, for null mean input sequences. The ACT can also be computed non-exactly for any input sequence, with low area complexity and low power consumption, utilizing the novel architecture described. However, as a trade-off, the ACT algorithm requires 10 non-uniformly sampled data points to calculate the eight-point DCT. This requirement can easily be satisfied for applications dealing with spatial signals such as image sensors and biomedical sensor arrays, by placing sensor elements in a non-uniform grid. In this work, a hardware architecture for the computation of the null mean ACT is proposed, followed by a novel architectures that extend the ACT for non-null mean signals. All circuits are physically implemented and tested using the Xilinx XC6VLX240T FPGA device and synthesized for 45 nm TSMC standard-cell library for performance assessment.
Nilanka T. Rajapaksha, Arjuna Madanayake, Renato J. Cintra, Jithra Adikari, Vassil S. Dimitrov
IEEE Trans. Computers3
2014 Low-complexity 8-point DCT approximations based on integer functions
Renato J. Cintra, Fábio M. Bayer, C. J. Tablada
Signal Process.1
2014 Analytic Expressions for Stochastic Distances Between Relaxed Complex Wishart Distributions
abstract
The scaled complex Wishart distribution is a widely used model for multilook full polarimetric synthetic aperture radar data whose adequacy is attested in this paper. Classification, segmentation, and image analysis techniques that depend on this model are devised, and many of them employ some type of dissimilarity measure. In this paper, we derive analytic expressions for four stochastic distances between relaxed scaled complex Wishart distributions in their most general form and in important particular cases. Using these distances, inequalities are obtained that lead to new ways of deriving the Bartlett and revised Wishart distances. The expressiveness of the four analytic distances is assessed with respect to the variation of parameters. Such distances are then used for deriving new tests statistics, which are proved to have asymptotic chi-square distribution. Adopting the test size as a comparison criterion, a sensitivity study is performed by means of Monte Carlo experiments suggesting that the Bhattacharyya statistic outperforms all the others. The power of the tests is also assessed. Applications to actual data illustrate the discrimination and homogeneity identification capabilities of these distances.
Alejandro C. Frery, Abraao D. C. Nascimento, Renato J. Cintra
IEEE Trans. Geosci. Remote. Sens.3
2014 Bias Correction and Modified Profile Likelihood Under the Wishart Complex Distribution
abstract
This paper proposes improved methods for the maximum likelihood (ML) estimation of the equivalent number of looks L. This parameter has a meaningful interpretation in the context of polarimetric synthetic aperture radar (PolSAR) images. Due to the presence of coherent illumination in their processing, PolSAR systems generate images which present a granular noise called speckle. As a potential solution for reducing such interference, the parameter L controls the signal-noise ratio. Thus, the proposal of efficient estimation methodologies for L has been sought. To that end, we consider first that a PolSAR image is well described by the scaled complex Wishart distribution. In recent years, Anfinsen have derived and analyzed estimation methods based on the ML and on trace statistical moments for obtaining the parameter L of the unscaled version of such probability law. This paper generalizes that approach. We present the second-order bias expression proposed by Cox and Snell for the ML estimator of this parameter. Moreover, the formula of the profile likelihood modified by Barndorff-Nielsen in terms of L is discussed. Such derivations yield two new ML estimators for the parameter L, which are compared to the estimators proposed by Anfinsen The performance of these estimators is assessed by means of Monte Carlo experiments, adopting three statistical measures as comparison criterion: the mean square error, the bias, and the coefficient of variation. Equivalently to the simulation study, an application to actual PolSAR data concludes that the proposed estimators outperform all the others in homogeneous scenarios.
Abraao D. C. Nascimento, Alejandro C. Frery, Renato J. Cintra
IEEE Trans. Geosci. Remote. Sens.3
2013 Parametric and nonparametric tests for speckled imagery
Renato J. Cintra, Alejandro C. Frery, Abraao D. C. Nascimento
Pattern Anal. Appl.1
2013 A Single-Channel Architecture for Algebraic Integer-Based 8 × 8 2-D DCT Computation
abstract
An area efficient row-parallel architecture is proposed for the real-time implementation of bivariate algebraic integer (AI) encoded 2-D discrete cosine transform (DCT) for image and video processing. The proposed architecture computes 8 × 8 2-D DCT transform based on the Arai DCT algorithm. An improved fast algorithm for AI-based 1-D DCT computation is proposed along with a single channel 2-D DCT architecture. The design improves on the four-channel AI DCT architecture that was published recently by reducing the number of integer channels to one and the number of eight-point 1-D DCT cores from five down to two. The architecture offers exact computation of 8 × 8 blocks of the 2-D DCT coefficients up to the FRS, which converts the coefficients from the AI representation to fixed-point format using the method of expansion factors. Prototype circuits corresponding to FRS blocks based on two expansion factors are realized, tested, and verified on FPGA-chip, using a Xilinx Virtex-6 XC6VLX240T device. Post place-and-route results show a 20% reduction in terms of area compared to the 2-D DCT architecture requiring five 1-D AI cores. The area-time and area-time2complexity metrics are also reduced by 23% and 22% respectively for designs with eight-bit input word length. The digital realizations are simulated up to place and route for ASICs using 45 nm CMOS standard cells. The maximum estimated clock rate is 951 MHz for the CMOS realizations indicating 7.608·109pixels/s and a 8 × 8 block rate of 118.875 MHz.
Amila Edirisuriya, Arjuna Madanayake, Renato J. Cintra, Vassil S. Dimitrov, Nilanka T. Rajapaksha
IEEE Trans. Circuits Syst. Video Technol.3
2013 Entropy-Based Statistical Analysis of PolSAR Data
abstract
Images obtained from coherent illumination processes are contaminated with speckle noise, with polarimetric synthetic aperture radar (PolSAR) imagery as a prominent example. With adequacy widely attested in the literature, the scaled complex Wishart distribution is an acceptable model for PolSAR data. In this perspective, we derive analytical expressions for the Shannon, Rényi, and restricted Tsallis entropy measurements under this model. Relationships between the derived measures and the parameters of the scaled Wishart law (i.e., the equivalent number of looks and the covariance matrix) are discussed. In addition, we obtain the asymptotic variances of the Shannon and Rényi entropy measurements when replacing distribution parameters by maximum-likelihood estimators. As a consequence, confidence intervals based on the Shannon and Rényi entropy measurements are also derived and proposed as new ways of capturing contrast. New hypothesis tests are additionally proposed using these results, and their performance is assessed using simulated and real data. In general terms, the test based on the Shannon entropy outperforms those based on Rényi entropy.
Alejandro C. Frery, Renato J. Cintra, Abraao D. C. Nascimento
IEEE Trans. Geosci. Remote. Sens.2
2012 Error-free VLSI architecture for the 2-D Daubechies 4-tap filter using algebraic integers
abstract
In this paper, a multi-encoding approach using wavelet based subband coding is proposed to accomplish error free calculations from exact representation of Daubechies 4-tap wavelet filter coefficients using the algebraic integer (AI) representation. By mapping the irrational coefficients to a convenient AI basis, the proposed architecture is designed employing a parallel channel model having two data paths carrying integer sequences. The computations done in the AI architecture are exactly accurate and are done entirely in a multiplier-free circuit. The design is implemented on a Xilinx Virtex-6 device at 172 MHz and hardware co-simulated with an ML605 board at 100 MHz. AI mapping facilitates simplicity and error-free calculations. The proposed architecture has a single final reconstruction step (FRS). Booth encoding has been adopted at the FRS to minimize the error which can be incurred at this point. This paper provides a hardware and power analysis for different bit lengths. An example output image sequence for the mandrill image is also provided.
Shiva Madishetty, Arjuna Madanayake, Renato J. Cintra, Dale H. Mugler, Vassil S. Dimitrov
ISCAS3
2012 A Row-Parallel 8 × 8 2-D DCT Architecture Using Algebraic Integer-Based Exact Computation
abstract
An algebraic integer (AI)-based time-multiplexed row-parallel architecture and two final reconstruction step (FRS) algorithms are proposed for the implementation of bivariate AI encoded 2-D discrete cosine transform (DCT). The architecture directly realizes an error-free 2-D DCT without using FRSs between row-column transforms, leading to an 8 × 8 2-D DCT that is entirely free of quantization errors in AI basis. As a result, the user-selectable accuracy for each of the coefficients in the FRS facilitates each of the 64 coefficients to have its precision set independently of others, avoiding the leakage of quantization noise between channels as is the case for published DCT designs. The proposed FRS uses two approaches based on: 1) optimized Dempster-Macleod multipliers, and 2) expansion factor scaling. This architecture enables low-noise high-dynamic range applications in digital video processing that requires full control of the finite-precision computation of the 2-D DCT. The proposed architectures and FRS techniques are experimentally verified and validated using hardware implementations that are physically realized and verified on field-programmable gate array (FPGA) chip. Six designs, for 4-bit and 8-bit input word sizes, using the two proposed FRS schemes, have been designed, simulated, physically implemented, and measured. The maximum clock rate and block rate achieved among 8-bit input designs are 307.787 MHz and 38.47 MHz, respectively, implying a pixel rate of 8 × 307.787≈2.462 GHz if eventually embedded in a real- time video-processing system. The equivalent frame rate is about 1187.35Hz for the image size of 1920 × 1080. All implementations are functional on a Xilinx Virtex-6 XC6VLX240T FPGA device.
Arjuna Madanayake, Renato J. Cintra, Denis Onen, Vassil S. Dimitrov, Nilanka T. Rajapaksha, Leonard T. Bruton, Amila Edirisuriya
IEEE Trans. Circuits Syst. Video Technol.2
2011 A new algorithm for double scalar multiplication over Koblitz curves
abstract
Koblitz curves are a special set of elliptic curves and have improved performance in computing scalar multiplication in elliptic curve cryptography due to the Frobenius endomorphism. Double-base number system approach for Frobenius expansion has improved the performance in single scalar multiplication. In this paper, we present a new algorithm to generate a sparse and joint τ-adic representation for a pair of scalars and its application in double scalar multiplication. The new algorithm is inspired from double-base number system. We achieve 12% improvement in speed against state-of-the-art τ-adic joint sparse form.
Jithra Adikari, Vassil S. Dimitrov, Renato J. Cintra
ISCAS3
2011 Algebraic integer based 8×8 2-D DCT architecture for digital video processing
abstract
A time-multiplexed row-parallel architecture is pro- posed for the real-time implementation of bivariate algebraic integer (AI) encoded 2-D discrete cosine transform (DCT) of images and video sequences. The architecture is based on the Arai algorithm with AI encoding. This leads to an 8×8 2-D DCT which is entirely free of quantization errors. The error free coefficients may be converted into a regular arithmetic format using a final reconstruction step (FRS) at the output stage. The accuracy of the FRS allows each of the 64 coefficients to have its precision set independent of other coefficients without the leakage of quantization noise between coefficient channels. Our architecture leads to low-noise applications in digital video compression, coding, and other image processing applications that rely on the fast systolic computation of the 2-D DCT. A prototype of the 2-D DCT is physically realized, tested, and verified on chip, using a Xilinx Virtex-4 S×35-10ff668 device. The maximum clock rate was Fclock= 121 MHz, implying an equivalent frame sample rate of 466 Hz, for an image frame size of 1920 × 1080, which is a common high definition video format.
Arjuna Madanayake, Renato J. Cintra, Denis Onen, Vassil S. Dimitrov, Leonard T. Bruton
ISCAS2
2011 A DCT Approximation for Image Compression
abstract
An orthogonal approximation for the 8-point discrete cosine transform (DCT) is introduced. The proposed transformation matrix contains only zeros and ones; multiplications and bit-shift operations are absent. Close spectral behavior relative to the DCT was adopted as design criterion. The proposed algorithm is superior to the signed discrete cosine transform. It could also outperform state-of-the-art algorithms in low and high image compression scenarios, exhibiting at the same time a comparable computational complexity.
Renato J. Cintra, Fábio M. Bayer
IEEE Signal Process. Lett.1
2010 Contrast in speckled imagery with stochastic distances
abstract
Synthetic aperture radar (SAR), ultrasound-B, laser, and sonar imagery are contaminated with speckle noise. The statistical modelling of such contamination is well described by the multiplicative model, which yields the G0distribution. In particular, reliable image contrast measures are sought in order to discriminate targets. To that end, we present statistical methods based on stochastic divergences and on the Kolmogorov-Smirnov distance for G0data. Their performance is quantified according to their test sizes and powers. A robustness analysis is also presented for several degrees of contamination. We show that the proposed tests based on triangular and arithmetic-geometric measures outperform the Kolmogorov-Smirnov distance.
Alejandro C. Frery, Abraao D. C. Nascimento, Renato J. Cintra
ICIP3
2010 Hypothesis Testing in Speckled Data With Stochastic Distances
abstract
Images obtained with coherent illumination, as is the case of sonar, ultrasound-B, laser, and synthetic aperture radar, are affected by speckle noise which reduces the ability to extract information from the data. Specialized techniques are required to deal with such imagery, which has been modeled by the${\cal G}^{0}$distribution and, under which, regions with different degrees of roughness and mean brightness can be characterized by two parameters; a third parameter, which is the number of looks, is related to the overall signal-to-noise ratio. Assessing distances between samples is an important step in image analysis; they provide grounds of the separability and, therefore, of the performance of classification procedures. This paper derives and compares eight stochastic distances and assesses the performance of hypothesis tests that employ them and maximum likelihood estimation. We conclude that tests based on the triangular distance have the closest empirical size to the theoretical one, while those based on the arithmetic–geometric distances have the best power. Since the power of tests based on the triangular distance is close to optimum, we conclude that the safest choice is using this distance for hypothesis testing, even when compared with classical distances as Kullback–Leibler and Bhattacharyya.
Abraao D. C. Nascimento, Renato J. Cintra, Alejandro C. Frery
IEEE Trans. Geosci. Remote. Sens.2
2009 Fragile watermarking using finite field trigonometrical transforms
Renato J. Cintra, Vassil S. Dimitrov, Hélio M. de Oliveira, Ricardo M. Campello de Souza
Signal Process. Image Commun.1
2004 Elliptic-cylindrical wavelets: the Mathieu wavelets
abstract
This note introduces a new family of wavelets and a multiresolution analysis that exploits the relationship between analyzing filters and Floquet's solution of Mathieu differential equations. The transfer function of both the detail and the smoothing filter is related to the solution of a Mathieu equation of the odd characteristic exponent. The number of notches of these filters can be easily designed. Wavelets derived by this method have potential application in the fields of optics and electromagnetism.
M. M. S. Lira, Hélio M. de Oliveira, Renato J. Cintra
IEEE Signal Process. Lett.3
2002 How to interpolate in arithmetic transform algorithms
abstract
In this paper, we propose a unified theory for arithmetic transform of a variety of discrete trigonometric transforms. The main contribution of this work is the elucidation of the interpolation process required in arithmetic transforms. We show that the interpolation method determines the transform to be computed. Several kernels were examined and asymptotic interpolation formulae were derived. Using the arithmetic transform theory, we also introduce a new algorithm for computing the discrete Hartley transform.
Renato J. Cintra, Hélio M. de Oliveira
ICASSP1