Daniel Llamocca

dblp:86/7940 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0003-1301-3655ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 56% Hardware accelerators and domain-specific architectures · 40% Interconnection networks and networks-on-chip · 5%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA implementation
0.522017
Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures · IEEE Trans. Image Process. 2017
Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform · IEEE Trans. Image Process. 2016
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator › convolution acceleration
convolution accelerator
0.312017
Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures · IEEE Trans. Image Process. 2017
Image and video processing › image reconstruction
reconstruction from projections
0.112016
Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform · IEEE Trans. Image Process. 2016
Interconnection networks and networks-on-chip › interconnect architecture
reconfigurable interconnect
0.112015
Field-Programmable Wiring Systems · Proc. IEEE 2015

Methods — techniques the papers use, named apart from their topics

singular value decomposition · 0.6discrete periodic radon transform · 0.6LU decomposition · 0.6fast transposition · 0.5circular shift register · 0.5block-based computation · 0.5
YearPublicationVenuePosition
2023 Fixed-point implementations for feed-forward artificial neural networks
Daniel Llamocca
Integr.1
2018 Fast and Parallel Computation of the Discrete Periodic Radon Transform on GPUs, Multicore CPUs and FPGAs
abstract
The Discrete Periodic Radon Transform (DPRT) has many important applications in reconstructing images from their projections and has recently been used in fast and scalable architectures for computing 2D convolutions. Unfortunately, the direct computation of the DPRT involves O(N3) additions and memory accesses that can be very costly in single-core architectures. The current paper presents new and efficient algorithms for computing the DPRT and its inverse on multi-core CPUs and GPUs. The results are compared against specialized hardware implementations (FPGAs/ASICs). The results provide significant evidence of the success of the new algorithms. On an 8-core CPU (Intel Xeon), with support for two threads per core, FastDirDPRT and FastDirInvDPRT achieve a speedup of approximately 10× (up to 12.83×) over the single-core CPU implementation. On a 2048-core GPU (GTX 980), FastRayDPRT and FastRayInvDPRT achieve speedups in the range of 526 (for 127 × 127) to 873 (for 1021 × 1021), which approximate ideal speedups of what can be achieved. The DPRT can be computed exactly and in real-time (30 frames per second) for 1471 × 1471 images using FastRayDPRT on the GPU. Furthermore, the GPU algorithms approximate the performance of an efficient FPGA implementation using 2N parallel cores at 100MHz.
Cesar Carranza, Marios S. Pattichis, Daniel Llamocca
ICIP3
2017 Self-reconfigurable architectures for HEVC Forward and Inverse Transform
Daniel Llamocca
J. Parallel Distributed Comput.1
2017 Fast 2D Convolutions and Cross-Correlations Using Scalable Architectures
abstract
The manuscript describes fast and scalable architectures and associated algorithms for computing convolutions and cross-correlations. The basic idea is to map 2D convolutions and cross-correlations to a collection of 1D convolutions and cross-correlations in the transform domain. This is accomplished through the use of the discrete periodic radon transform for general kernels and the use of singular value decomposition -LU decompositions for low-rank kernels. The approach uses scalable architectures that can be fitted into modern FPGA and Zynq-SOC devices. Based on different types of available resources, for P × P blocks, 2D convolutions and cross-correlations can be computed in just O(P) clock cycles up to O(P2) clock cycles. Thus, there is a trade-off between performance and required numbers and types of resources. We provide implementations of the proposed architectures using modern programmable devices (Virtex-7 and Zynq-SOC). Based on the amounts and types of required resources, we show that the proposed approaches significantly outperform current methods.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
IEEE Trans. Image Process.2
2016 Fast and Scalable Computation of the Forward and Inverse Discrete Periodic Radon Transform
abstract
The discrete periodic radon transform (DPRT) has extensively been used in applications that involve image reconstructions from projections. Beyond classic applications, the DPRT can also be used to compute fast convolutions that avoids the use of floating-point arithmetic associated with the use of the fast Fourier transform. Unfortunately, the use of the DPRT has been limited by the need to compute a large number of additions and the need for a large number of memory accesses. This paper introduces a fast and scalable approach for computing the forward and inverse DPRT that is based on the use of: a parallel array of fixed-point adder trees; circular shift registers to remove the need for accessing external memory components when selecting the input data for the adder trees; an image block-based approach to DPRT computation that can fit the proposed architecture to available resources; and fast transpositions that are computed in one or a few clock cycles that do not depend on the size of the input image. As a result, for an N × N image (N prime), the proposed approach can compute up to N(2) additions per clock cycle. Compared with the previous approaches, the scalable approach provides the fastest known implementations for different amounts of computational resources. For example, for a 251×251 image, for approximately 25% fewer flip-flops than required for a systolic implementation, we have that the scalable DPRT is computed 36 times faster. For the fastest case, we introduce optimized just 2N + ⌈log(2) N⌉ + 1 and 2N + 3 ⌈log(2) N⌉ + B + 2 cycles, architectures that can compute the DPRT and its inverse in respectively, where B is the number of bits used to represent each input pixel. On the other hand, the scalable DPRT approach requires more 1-b additions than for the systolic implementation and provides a tradeoff between speed and additional 1-b additions. All of the proposed DPRT architectures were implemented in VHSIC Hardware Description Language (VHDL) and validated using an Field-Programmable Gate Array (FPGA) implementation.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
IEEE Trans. Image Process.2
2015 A scalable pipelined architecture for biomimetic vision sensors
abstract
This work presents a novel scalable and fully pipelined digital hardware implementation for a biomimetic vision sensor that mimics the compound eye of the common house fly. The hardware includes IIR filters, FIR filters, and arithmetic units. By employing techniques such as scattered look-ahead decomposition, retiming, and distributed arithmetic, a fully parallel and pipelined system that produces new outputs at every clock cycle is achieved. Trade-offs among design parameters, accuracy, and resource usage are explored. A generic and parameterized architecture is developed in VHDL and validated using an FPGA implementation. The developed hardware shows great promise as it allows for the implementation of fly-inspired sensing and control algorithms whose motion and edge detection capabilities surpass those of traditional image processing systems.
Daniel Llamocca, Brian K. Dean
FPL1
2015 Field-Programmable Wiring Systems
abstract
Field-programmable wiring systems refer to methods and hardware that can maintain the interconnection of components of different types. Generally, field-programmable wiring systems support the use of multidomain fabrics that can be used to route analog, power, digital signals, optical, microwave signals, etc. This paper reviews fundamental concepts associated with the practical implementation of field-programmable wiring systems. The paper also provides different implementation examples and discusses a list of challenges and recommendations for future work in this area.
Víctor Murray, Marios S. Pattichis, Daniel Llamocca, James Lyke
Proc. IEEE3
2015 Dynamic Energy, Performance, and Accuracy Optimization and Management Using Automatically Generated Constraints for Separable 2D FIR Filtering for Digital Video Processing
abstract
There is strong interest in the development of dynamically reconfigurable systems that can meet real-time constraints on energy, performance, and accuracy. The generation of real-time constraints will significantly expand the applicability of dynamically reconfigurable systems to new domains, such as digital video processing. We develop a dynamically reconfigurable 2D FIR filtering system that can meet real-time constraints in energy, performance, and accuracy (EPA). The real-time constraints are automatically generated based on user input, image types associated with video communications, and video content. We first generate a set of Pareto-optimal realizations, described by their EPA values and associated 2D FIR hardware description bitstreams. Dynamic management is then achieved by selecting Pareto-optimal realizations that meet the automatically generated time-varying EPA constraints. We validate our approach using three different 2D Gaussian filters. Filter realizations are evaluated in terms of the required energy per frame, accuracy of the resulting image, and performance in frames per second. We demonstrate dynamic EPA management by applying a Difference of Gaussians (DOG) filter to standard video sequences. For video frame sizes that are equal to or larger than the VGA resolution, compared to a static implementation, our dynamic system provides significant reduction in the total energy consumption (>30%).
Daniel Llamocca, Marios S. Pattichis
ACM Trans. Reconfigurable Technol. Syst.1
2014 A scalable architecture for implementing the fast discrete periodic radon transform for prime sized images
abstract
The Discrete Periodic Radon Transform (DPRT) has many important applications in image processing that are associated with reconstructing objects from projections (e.g., computed tomography [1]) or image restoration (e.g., [2]). Thus, there is strong interest in the development of fast algorithms and architectures for computing the DPRT. This paper introduces a scalable hardware architecture and associated algorithm for computing the DPRT for prime-sized images. For square images of size N × N, N prime, the DPRT requires N2(N - 1) additions for calculating image projections along a minimal number of prime directions. The proposed approach can compute the DPRT in [N/2h] N + 2N + h clock cycles, h = 1, ..., [log2N], where h is a scaling factor that is used to control the required hardware resources that are needed to implement the fast DPRT. Compared to previous approaches, a fundamental contribution of the proposed architecture is that it allows effective implementations based on different constraints on the resources.
Cesar Carranza, Daniel Llamocca, Marios S. Pattichis
ICIP2
2013 A Dynamically Reconfigurable Pixel Processor System Based on Power/Energy-Performance-Accuracy Optimization
abstract
We introduce a dynamically reconfigurable framework for implementing single-pixel operations. The system relies on a multiobjective optimization scheme that generates Pareto-optimal realizations in the power/energy-performance-accuracy (PPA/EPA) spaces. The Pareto-optimal realizations and their PPA/EPA values are stored in DDR-SDRAM and can be chosen dynamically to meet time-varying constraints. Results are shown in terms of power, accuracy (peak signal-to-noise ratio) of the resulting image, and performance in frames per second. Dynamic PPA/EPA management is implemented using dynamic partial reconfiguration and dynamic frequency control.
Daniel Llamocca, Marios S. Pattichis
IEEE Trans. Circuits Syst. Video Technol.1
2012 Dynamic multiobjective optimization management of the Energy-Performance-Accuracy space for separable 2-D complex filters
abstract
We present a dynamic framework for 2D complex filter implementation that is based on a multi-objective optimization scheme that generates Pareto-optimal realizations from the Energy-Performance-Accuracy (EPA) space. The EPA space is created by evaluating the 2D complex filter realizations in terms of their required energy, accuracy, and performance. Dynamic EPA management, carried out via Dynamic Partial Reconfiguration (DPR) and Dynamic Frequency Control, then consists on selecting Pareto-optimal realizations that meet time-varying EPA requirements. We demonstrate dynamic EPA management by applying a complex filter to a standard video sequence.
Daniel Llamocca, Cesar Carranza, Marios S. Pattichis
FPL1
2011 Separable FIR Filtering in FPGA and GPU Implementations: Energy, Performance, and Accuracy Considerations
abstract
Digital video processing requires significant hardware resources to achieve acceptable performance. Digital video processing based on dynamic partial reconfiguration (DPR) allows the designers to control resources based on energy, performance, and accuracy considerations. In this paper, we present a dynamically reconfigurable implementation of a 2D FIR filter where the number of coefficients and coefficients values can be varied to control energy, performance, and precision requirements. We also present a high-performance GPU implementation to help understand the trade-offs between these two technologies. Results using a standard example of 2D Difference of Gaussians (DOG) filter indicate that the DPR implementation can deliver real-time performance with energy per frame consumption that is an order of magnitude less than the GPU. On the other hand, at significantly higher energy consumption levels, the GPU implementation can deliver very high performance.
Daniel Llamocca, Cesar Carranza, Marios S. Pattichis
FPL1
2010 Using support vector machines for anomalous change detection
abstract
We cast anomalous change detection as a binary classification problem, and use a support vector machine (SVM) to build a detector that does not depend on assumptions about the underlying data distribution. To speed up the computation, our SVM is implemented, in part, on a graphical processing unit. Results on real and simulated anomalous changes are used to compare performance to algorithms which effectively assume a Gaussian distribution.
Ingo Steinwart, James Theiler, Daniel Llamocca
IGARSS3
2009 A dynamically reconfigurable parallel pixel processing system
abstract
We describe a dynamically reconfigurable image processing system that reaches real time video processing performances despite reconfiguration time overhead. The system is composed of reconfigurable pixel processing units set to process several pixels in parallel. We present a scheme for optimizing a LUT-based architecture by directly mapping it into the Xilinx FPGA CLB primitives. Internally controlled dynamic partial reconfiguration is to modify the LUT values at run-time without stalling the overall operation. The combination of optimized implementations with CLB primitives and dynamic partial reconfiguration leads to multifunctional, area-efficient, and highperformance realizations of LUT-based pixel processing systems. We present results from a dynamically reconfigurable high-performance LUT-based image/video processing system. Experimental measurements show that the system achieves speeds of 226 Mbps with small resource utilization. The architecture can dynamically reconfigure the image processing operation at each new frame and still reach real-time video processing speeds (640 times 480 graylevel frames). We also evaluate the effect that increasing partial reconfiguration rates have on the system's overall performance.
Daniel Llamocca, Marios S. Pattichis, G. Alonzo Vera
FPL1