Cesar Carranza

dblp:49/10313 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0003-1222-0118ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-authorSystems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 50% Reconfigurable computing and FPGAs · 50%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA implementation
0.522017
Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures · IEEE Trans. Image Process. 2017
Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform · IEEE Trans. Image Process. 2016
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator › convolution acceleration
convolution accelerator
0.312017
Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures · IEEE Trans. Image Process. 2017
Image and video processing › image reconstruction
reconstruction from projections
0.112016
Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform · IEEE Trans. Image Process. 2016

Methods — techniques the papers use, named apart from their topics

singular value decomposition · 0.6discrete periodic radon transform · 0.6LU decomposition · 0.6fast transposition · 0.5circular shift register · 0.5block-based computation · 0.5
YearPublicationVenuePosition
2018 Fast and Parallel Computation of the Discrete Periodic Radon Transform on GPUs, Multicore CPUs and FPGAs
abstract
The Discrete Periodic Radon Transform (DPRT) has many important applications in reconstructing images from their projections and has recently been used in fast and scalable architectures for computing 2D convolutions. Unfortunately, the direct computation of the DPRT involves O(N3) additions and memory accesses that can be very costly in single-core architectures. The current paper presents new and efficient algorithms for computing the DPRT and its inverse on multi-core CPUs and GPUs. The results are compared against specialized hardware implementations (FPGAs/ASICs). The results provide significant evidence of the success of the new algorithms. On an 8-core CPU (Intel Xeon), with support for two threads per core, FastDirDPRT and FastDirInvDPRT achieve a speedup of approximately 10× (up to 12.83×) over the single-core CPU implementation. On a 2048-core GPU (GTX 980), FastRayDPRT and FastRayInvDPRT achieve speedups in the range of 526 (for 127 × 127) to 873 (for 1021 × 1021), which approximate ideal speedups of what can be achieved. The DPRT can be computed exactly and in real-time (30 frames per second) for 1471 × 1471 images using FastRayDPRT on the GPU. Furthermore, the GPU algorithms approximate the performance of an efficient FPGA implementation using 2N parallel cores at 100MHz.
Cesar Carranza, Marios S. Pattichis, Daniel Llamocca
ICIP1
2017 Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures
abstract
The manuscript describes fast and scalable architectures and associated algorithms for computing convolutions and cross-correlations. The basic idea is to map 2D convolutions and cross-correlations to a collection of 1D convolutions and cross-correlations in the transform domain. This is accomplished through the use of the discrete periodic radon transform for general kernels and the use of singular value decomposition -LU decompositions for low-rank kernels. The approach uses scalable architectures that can be fitted into modern FPGA and Zynq-SOC devices. Based on different types of available resources, for P × P blocks, 2D convolutions and cross-correlations can be computed in just O(P) clock cycles up to O(P2) clock cycles. Thus, there is a trade-off between performance and required numbers and types of resources. We provide implementations of the proposed architectures using modern programmable devices (Virtex-7 and Zynq-SOC). Based on the amounts and types of required resources, we show that the proposed approaches significantly outperform current methods.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
IEEE Trans. Image Process.1
2016 Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform
abstract
The discrete periodic radon transform (DPRT) has extensively been used in applications that involve image reconstructions from projections. Beyond classic applications, the DPRT can also be used to compute fast convolutions that avoids the use of floating-point arithmetic associated with the use of the fast Fourier transform. Unfortunately, the use of the DPRT has been limited by the need to compute a large number of additions and the need for a large number of memory accesses. This paper introduces a fast and scalable approach for computing the forward and inverse DPRT that is based on the use of: a parallel array of fixed-point adder trees; circular shift registers to remove the need for accessing external memory components when selecting the input data for the adder trees; an image block-based approach to DPRT computation that can fit the proposed architecture to available resources; and fast transpositions that are computed in one or a few clock cycles that do not depend on the size of the input image. As a result, for an N × N image (N prime), the proposed approach can compute up to N(2) additions per clock cycle. Compared with the previous approaches, the scalable approach provides the fastest known implementations for different amounts of computational resources. For example, for a 251×251 image, for approximately 25% fewer flip-flops than required for a systolic implementation, we have that the scalable DPRT is computed 36 times faster. For the fastest case, we introduce optimized just 2N + ⌈log(2) N⌉ + 1 and 2N + 3 ⌈log(2) N⌉ + B + 2 cycles, architectures that can compute the DPRT and its inverse in respectively, where B is the number of bits used to represent each input pixel. On the other hand, the scalable DPRT approach requires more 1-b additions than for the systolic implementation and provides a tradeoff between speed and additional 1-b additions. All of the proposed DPRT architectures were implemented in VHSIC Hardware Description Language (VHDL) and validated using an Field-Programmable Gate Array (FPGA) implementation.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
IEEE Trans. Image Process.1
2014 A scalable architecture for implementing the fast discrete periodic radon transform for prime sized images
abstract
The Discrete Periodic Radon Transform (DPRT) has many important applications in image processing that are associated with reconstructing objects from projections (e.g., computed tomography [1]) or image restoration (e.g., [2]). Thus, there is strong interest in the development of fast algorithms and architectures for computing the DPRT. This paper introduces a scalable hardware architecture and associated algorithm for computing the DPRT for prime-sized images. For square images of size N × N, N prime, the DPRT requires N2(N - 1) additions for calculating image projections along a minimal number of prime directions. The proposed approach can compute the DPRT in [N/2h] N + 2N + h clock cycles, h = 1, ..., [log2N], where h is a scaling factor that is used to control the required hardware resources that are needed to implement the fast DPRT. Compared to previous approaches, a fundamental contribution of the proposed architecture is that it allows effective implementations based on different constraints on the resources.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
ICIP1
2012 Dynamic multiobjective optimization management of the Energy-Performance-Accuracy space for separable 2-D complex filters
abstract
We present a dynamic framework for 2D complex filter implementation that is based on a multi-objective optimization scheme that generates Pareto-optimal realizations from the Energy-Performance-Accuracy (EPA) space. The EPA space is created by evaluating the 2D complex filter realizations in terms of their required energy, accuracy, and performance. Dynamic EPA management, carried out via Dynamic Partial Reconfiguration (DPR) and Dynamic Frequency Control, then consists on selecting Pareto-optimal realizations that meet time-varying EPA requirements. We demonstrate dynamic EPA management by applying a complex filter to a standard video sequence.
Daniel Llamocca, Cesar Carranza, Marios S. Pattichis
FPL2
2011 Separable FIR Filtering in FPGA and GPU Implementations: Energy, Performance, and Accuracy Considerations
abstract
Digital video processing requires significant hardware resources to achieve acceptable performance. Digital video processing based on dynamic partial reconfiguration (DPR) allows the designers to control resources based on energy, performance, and accuracy considerations. In this paper, we present a dynamically reconfigurable implementation of a 2D FIR filter where the number of coefficients and coefficients values can be varied to control energy, performance, and precision requirements. We also present a high-performance GPU implementation to help understand the trade-offs between these two technologies. Results using a standard example of 2D Difference of Gaussians (DOG) filter indicate that the DPR implementation can deliver real-time performance with energy per frame consumption that is an order of magnitude less than the GPU. On the other hand, at significantly higher energy consumption levels, the GPU implementation can deliver very high performance.
Daniel Llamocca, Cesar Carranza, Marios S. Pattichis
FPL2