Francisco Argüello

dblp:04/501 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-9279-5426ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorArtificial intelligence and machine learning · 2Computer networks · 1
YearPublicationVenuePosition
2026 Distributed multi-GPU algorithm for accurate registration of UAV-based multispectral and multitemporal orthomosaics
abstract
Abstract Accurate registration of high-resolution multispectral UAV orthomosaics acquired on different dates and sensors is essential for a wide range of remote sensing applications. However, this task remains challenging due to variations in acquisition conditions, including seasonal changes, differences in illumination, weather, and sensor characteristics. This article presents a parallel multilevel registration method that combines and improves two existing algorithms: HSI-KAZE, a feature-based approach, and HYFM, an area-based method. The three-level approach first applies an optimized HSI-KAZE (OHSI-KAZE) for coarse estimation of scale, rotation, and translation, followed by HYFM for fine correction. The multi-node multi-GPU proposed implementation, leveraging MPI, OpenMP, and CUDA, enables efficient processing on GPU-accelerated HPC clusters. Experiments on six real multispectral orthomosaic pairs from river environments, compared against state-of-the-art classical and deep learning methods, achieve high registration accuracy with RMSE values below 1.57 pixels and a 30 $$\times $$ × speedup, confirming both the accuracy and scalability of the proposed method.
Daniel Fuentes, Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello
J. Supercomput.4
2025 Attention-Based Convolutional Neural Network for Anomaly Detection in Multispectral Images of Semi-Natural Ecosystems
abstract
The monitoring of semi-natural ecosystems has become increasingly critical due to the rising impact of ecological disturbances, including natural disasters and unauthorized human-made constructions. Anomaly detection (AD) in multispectral imagery serves as a fundamental tool in this context. Deep learning-based techniques are particularly effective at capturing the intricate spectral and spatial patterns of anomalies. This paper proposes a new AD technique called ACNN, designed to enhance AD performance in multispectral images of high spatial resolution for the detection of human-made constructions. The model integrates attention mechanisms to prioritize informative features while suppressing irrelevant background information, thereby improving sensitivity to subtle and rare anomalies. Experimental results on multispectral datasets from semi-natural ecosystems show that the proposed approach outperforms existing deep learning (DL) techniques in terms of detection accuracy. These findings highlight the potential of attention-based models as a robust framework for environmental monitoring and AD in complex remote sensing scenarios.
Javier Lopez-Fandino, Álvaro Ordóñez, Pablo Quesada-Barriuso, Alberto S. Garea, Francisco Argüello, Dora Blanco Heras
IEEE Geosci. Remote. Sens. Lett.5
2024 Region-Based Multispectral Image Registration on Heterogeneous Computing Platforms
abstract
Feature-based methods are widely used for the registration of remote sensing images because of their robustness to viewpoint, scale, and light changes. However, they are computationally demanding, especially when dealing with multi or hyperspectral images. Hyperspectral Maximally Stable Extremal Regions (HSI-MSER) is a hyperspectral remote sensing image registration method based on MSER for feature detection and Scale Invariant Feature Transform (SIFT) for feature description. This article presents a first approach to a parallel implementation of the HSI-MSER algorithm for the registration of multispectral images on a heterogeneous computing platform. The results of the registration capabilities under extreme scaling and rotating conditions show that the proposed parallel implementation obtains a speedup of 3.33× compared to the sequential implementation making it suitable for applications with execution time constraints.
Daniel Del Castillo, Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello
IGARSS4
2024 Using heterogeneous computing and edge computing to accelerate anomaly detection in remotely sensed multispectral images
abstract
Abstract This paper proposes a parallel algorithm exploiting heterogeneous computing and edge computing for anomaly detection (AD) in remotely sensed multispectral images. These images present high spatial resolution and are captured onboard unmanned aerial vehicles. AD is applied to identify patterns within an image that do not conform to the expected behavior. In this paper, the anomalies correspond to human-made constructions that trigger alarms related to the integrity of fluvial ecosystems. An algorithm based on extracting spatial information by using extinction profiles (EPs) and detecting anomalies by using the Reed–Xiaoli (RX) technique is proposed. The parallel algorithm presented in this paper is designed to be executed on multi-node heterogeneous computing platforms that include nodes with multi-core central processing units (CPUs) and graphics processing units (GPUs) and on a mobile embedded system consisting of a multi-core CPU and a GPU. The experiments are carried out on nodes of the FinisTerrae III supercomputer and, with the objective of analyzing its efficiency under different energy consumption scenarios, on a Jetson AGX Orin.
Javier Lopez-Fandino, Dora Blanco Heras, Francisco Argüello
J. Supercomput.3
2023 Prospective Comparison of SURF and Binary Keypoint Descriptors for Fast Hyperspectral Remote Sensing Registration
abstract
Image registration is a crucial process that involves determining the geometric transformation required to align multiple images. It plays a vital role in various remote sensing image processing tasks that involve analyzing changes among images. To enable real-time response, it is essential to have computationally efficient registration algorithms, especially when dealing with large datasets as is the case of hyperspectral images. This article presents a comparative analysis of two descriptors used to characterize local features of images prior to their matching and registration. The objective is to analyze whether the LATCH binary keypoint descriptor, which produces compact descriptors, provides similar results to the gradient-based SURF descriptor in terms of execution time and registration precision. To obtain the best computational performance, multithreaded implementations using OpenMP have been proposed. LATCH has proven to be 7× faster and as reliable as SURF in terms of accuracy on scale differences of up to 1.2×.
Adrián Rodríguez-Molina, Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello, José F. López
IGARSS4
2023 A hybrid CUDA, OpenMP, and MPI parallel TCA-based domain adaptation for classification of very high-resolution remote sensing images
abstract
Abstract Domain Adaptation (DA) is a technique that aims at extracting information from a labeled remote sensing image to allow classifying a different image obtained by the same sensor but at a different geographical location. This is a very complex problem from the computational point of view, specially due to the very high-resolution of multispectral images. TCANet is a deep learning neural network for DA classification problems that has been proven as very accurate for solving them. TCANet consists of several stages based on the application of convolutional filters obtained through Transfer Component Analysis (TCA) computed over the input images. It does not require backpropagation training, in contrast to the usual CNN-based networks, as the convolutional filters are directly computed based on the TCA transform applied over the training samples. In this paper, a hybrid parallel TCA-based domain adaptation technique for solving the classification of very high-resolution multispectral images is presented. It is designed for efficient execution on a multi-node computer by using Message Passing Interface (MPI), exploiting the available Graphical Processing Units (GPUs), and making efficient use of each multicore node by using Open Multi-Processing (OpenMP). As a result, an accurate DA technique from the point of view of classification and with high speedup values over the sequential version is obtained, increasing the applicability of the technique to real problems.
Alberto S. Garea, Dora Blanco Heras, Francisco Argüello, Begüm Demir
J. Supercomput.3
2022 Multi-GPU Registration of High-Resolution Multispectral Images Using HSI-KAZE in a Cluster System
abstract
Feature-based registration methods have been demonstrated to be very effective to register multispectral images with large distortions or transformations despite the higher execution time that they require.In this paper, a first approach to a multi-node, multi-GPU implementation of the Hyperspectral KAZE (HSI-KAZE) method for co-registering bands and multispectral images is presented.Different multispectral datasets are distributed among the available nodes of a cluster using MPI and exploiting the parallel stream-based capabilities of the GPUs inside each node using CUDA.
Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello
IGARSS3
2021 Comparing Area-Based and Feature-Based Methods for Co-Registration of Multispectral Bands on GPU
abstract
Registration is required as a previous step for processing multispectral images. The different bands captured by each sensor for each image, as well as the different images corresponding to the same area, need to be aligned. In this paper, a 2-level registration scheme comparing the results obtained by the hyperspectral Fourier-Mellin (HYFM) and hyperspectral KAZE (HSI-KAZE) registration methods is proposed. It is designed for efficient implementation in a multi-GPU system in which different scenes are registered in parallel on different GPUs.
Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello
IGARSS3
2021 GPU accelerated waterpixel algorithm for superpixel segmentation of hyperspectral images
Pablo Quesada-Barriuso, Dora Blanco Heras, Francisco Argüello
J. Supercomput.3
2020 GPU-accelerated registration of hyperspectral images using KAZE features
Álvaro Ordóñez, Francisco Argüello, Dora Blanco Heras, Begüm Demir
J. Supercomput.2
2019 Surf-Based Registration for Hyperspectral Images
abstract
The alignment of images, also known as registration, is a relevant task in the processing of hyperspectral images. Among the feature-based registration methods, Speeded Up Robust Features (SURF) has been proposed as a computationally efficient approach. In this paper HSI-SURF is proposed. This is a method to register hyperspectral remote sensing images based on SURF that takes advantage of the full spectral information of the images. In this sense, the proposed method selects specific bands of the images and adapts the keypoint descriptor and the matching stages to benefit from the spectral information, thus increasing the effectiveness of the registration.
Álvaro Ordóñez, Dora Blanco Heras, Francisco Argüello
IGARSS3
2019 Extended attribute profiles on GPU applied to hyperspectral image classification
Pedro G. Bascoy, Pablo Quesada-Barriuso, Dora Blanco Heras, Francisco Argüello, Begüm Demir, Lorenzo Bruzzone
J. Supercomput.4
2019 Caffe CNN-based classification of hyperspectral images on GPU
Alberto S. Garea, Dora Blanco Heras, Francisco Argüello
J. Supercomput.3
2018 Stacked Autoencoders for Multiclass Change Detection in Hyperspectral Images
abstract
Change detection (CD) in multitemporal datasets is a key task in remote sensing. In this paper, a scheme to perform multiclass CD for remote sensing hyperspectral datasets extracting features by means of Stacked Autoencoders (SAEs) is introduced. The scheme combines multiclass and binary CD to obtain an accurate multiclass change map. The multiclass CD begins with the fusion of the multitemporal data followed by Feature Extraction (FE) by SAEs. The binary CD is based on the spectral information by calculating pixel-wise distances and thresholding, and it also incorporates spatial information through watershed segmentation. The processed image is filtered by using the binary CD map and later classified by a Support Vector Machine or an Extreme Learning Machine algorithm. The scheme was evaluated over a multitemporal hyperspectral dataset obtained from the Hyperion sensor. Experimental results show the effectiveness of the proposed scheme using a SAE for extracting the relevant features of the fused information when compared to other published FE methods.
Javier Lopez-Fandino, Alberto S. Garea, Dora Blanco Heras, Francisco Argüello
IGARSS4
2016 Evolutionary cellular automata based approach to high-dimensional image segmentation for GPU projection
abstract
This paper proposes an intrinsically distributed cellular automata (CA) based approach to address the perennial problem of real time segmentation and classification of high dimensional images, such as remote sensing hyperspectral images. This approach is efficiently implemented on GPUs providing results that improve on the state of the art algorithms presented in the literature. It is based on the evolutionary generation of the CA rule sets under two basic premises: During the segmentation process, the CAs must work over the whole dimensionality of the images without any projection onto lower dimensionalities, and the rule sets that are generated must be adapted to the segmentation level required by the user. The performance of the approach is tested over a benchmark set of well-known hyperspectral images and the results compared to the state of the art in the literature for two implementations, one using a SVM based classification stage and another that considers an ELM based classification stage.
Becerra Priego, Richard J. Duro, Javier Lopez-Fandino, Dora Blanco Heras, Francisco Argüello
IJCNN5
2012 Memory Hierarchy Optimization for Large Tridiagonal System Solvers on GPU
abstract
Nowadays GPUs are commodity hardware containing hundreds of cores and supporting thousands of threads that can be used to accelerate a wide range of applications. From a programmer's perspective, GPUs offer a stream processing model which requires the application of new techniques to exploit their capabilities. In this paper we present the application of the split-and-merge technique to the following parallel tridiagonal system solvers on the GPU: cyclic reduction and recursive doubling. The split-and-merge technique naturally splits the algorithm flow in parallel paths that can be solved in shared memory, and later merged in global memory. In this way, we can solve large systems of equations efficiently exploiting the memory hierarchy of the GPU. The results obtained show a significant acceleration compared with the direct implementation of the algorithms on the GPU.
Julián Lamas-Rodríguez, Francisco Argüello, Dora Blanco Heras, Montserrat Bóo
ISPA2
2012 Efficient GPU Asynchronous Implementation of a Watershed Algorithm Based on Cellular Automata
abstract
The watershed transform is a widely used method for non-supervised image segmentation, especially suitable for low-contrast images. In this paper we show that an algorithm calculating the watershed transform based on a cellular automaton is a good choice for the most recent GPU architectures, especially when the synchronization rules are relaxed. In particular we compare a synchronous and an asynchronous implementation of the algorithm. The results show high speedups for both implementations, especially for the asynchronous one, indicating the potential of this kind of algorithms for new architectures based on hundreds of cores.
Pablo Quesada-Barriuso, Dora Blanco Heras, Francisco Argüello
ISPA3
2012 Efficient segmentation of hyperspectral images on commodity GPUs
abstract
The techniques for segmentation and classification of hyperspectral images are very costly, which makes them good candidates for parallel and, in particular, GPU processing. In this paper we present a GPU implementation of a segmentation strategy for hyperspectral images consisting in the calculation of a morphological gradient operator that reduces the dimensionality of the hyperspectral image followed by the calculation of a watershed transform over the resulting 2D image. We have studied the main issues for the efficient implementation of the algorithms in GPU: the exploitation of thousands of threads available in this architecture and the adequate use of the device bandwidth. The tests show the efficiency of the GPU implementation indicating that the processing of hyperspectral images can be performed in real-time even on commodity GPUs like the one used in the experiments.
Pablo Quesada-Barriuso, Francisco Argüello, Dora Blanco Heras
KES2
2012 The split-and-merge method in general purpose computation on GPUs
Francisco Argüello, Dora Blanco Heras, Montserrat Bóo, Julián Lamas-Rodríguez
Parallel Comput.1
2002 Architecture for wavelet packet transform based on lifting steps
Francisco Argüello, Juan López, María A. Trenas, Emilio L. Zapata
Parallel Comput.1
2001 A Data-Parallel Formulation for Divide and Conquer Algorithms
abstract
This paper presents a general data-parallel formulation for a class of problems based on the divide and conquer strategy. A combination of three techniques—mapping vectors, index-digit permutations and space-filling curves—are used to reorganize the algorithmic dataflow, providing great flexibility to efficiently exploit data locality and to reduce and optimize communications. In addition, these techniques allow the easy translation of the reorganized dataflows into HPF (High Performance Fortran) constructs. Finally, experimental results on the Cray T3E validate our method.
Margarita Amor, Francisco Argüello, Juan López, Oscar G. Plata, Emilio L. Zapata
Comput. J.2
2000 Architecture for Wavelet Packet Transform with Best Tree Searching
abstract
Wavelet Packet Transform (WPT) provides good spectral and temporal resolutions in arbitrary regions of the time-frequency plane. Given an additive cost function, a best-tree searching algorithm allows the selection of the best basis for a given signal according to this function. This adaptive choice of the time-frequency tiling benefits most of the applications where the standard Wavelet Transform (WT) has already shown to be useful. Though many specific architectures have been proposed in the literature for the WT, it is not the case for WPT. In this work we present a specific architecture for WPT which implements the best-tree searching algorithm.
María A. Trenas, Juan López, Manuel Sánchez, Emilio L. Zapata, Francisco Argüello
ASAP5
2000 An architecture for wavelet-packet based speech enhancement for hearing aids
abstract
Wavelet packets have been applied in order to compensate the speech signal to improve the intelligibility for a common hearing impairment known as recruitment of loudness, a sensorineural hearing loss of cochlear origin. We present an architecture that allows selection of the best decomposition tree for each patient, in order to apply this wavelet-packet based parametric compression algorithm.
María A. Trenas, Juan López, Emilio L. Zapata, Francisco Argüello
ICASSP4
1998 A memory system supporting the efficient SIMD computation of the two dimensional DWT
abstract
Real time image processing uses SIMD engines to accelerate the computation of algorithms such as the DCT, FFT or DWT. So, a good skewing scheme becomes essential for avoiding memory bank conflicts. A memory system is introduced for the efficient in-place computation of such transforms. It consists of M=2/sup m/ memory modules, providing parallel access to M image points whose patterns are a row or a column, the interval in both cases being 2/sup l/, l/spl ges/0. The efficiency of our design is proved through the computation of the 2D DWT.
María A. Trenas, Juan López, Emilio L. Zapata, Francisco Argüello
ICASSP4
1997 Unified Framework for the Parallelization of Divide and Conquer Based Tridiagonal Systems
Juan López, Oscar G. Plata, Francisco Argüello, Emilio L. Zapata
Parallel Comput.3
1997 High-performance VLSI architecture for the Viterbi algorithm
abstract
The Viterbi (1967) algorithm (VA) is known to be an efficient method for the realization of maximum-likelihood (ML) decoding of convolutional codes. The VA is characterized by a graph, called a trellis, which defines the transitions between states. To define an area efficient architecture for the VA is equivalent to obtaining an efficient mapping of the trellis. We present a methodology that permits the efficient hardware mapping of the VA onto a processor network of arbitrary size. This formal model is employed for the partitioning of the computations among an arbitrary number of processors in such a way that the data are recirculated, optimizing the use of the PEs and the communications. Therefore, the algorithm is mapped onto a column of processing elements and an optimal design solution is obtained for a particular set of area and/or speed constraints. Furthermore, the management of the surviving path memory for its mapping and distribution among the processors was studied. As a result, we obtain a regular and modular design appropriate for its VLSI implementation in which the only necessary communications between processors are the data recirculations between stages.
Montserrat Bóo, Francisco Argüello, Javier D. Bruguera, Ramón Doallo, Emilio L. Zapata
IEEE Trans. Commun.2
1996 High-Speed Viterbi Decoder: An Efficient Scheduling Method to Exploit the Pipelining
abstract
The main part of the Viterbi algorithm is a nonlinear feedback loop which presents a bottleneck for high-speed implementations. We present a novel scheduling scheme that allows increasing the available speed of the system. This is done through the utilization of look-ahead techniques to compute non-sequential data and, in this way, break the recursivity of the algorithm. This permits introducing pipelining. As a result, we obtain a speed growth comparable to previous parallel solutions, but with less hardware cost.
Montserrat Bóo, Francisco Argüello, Javier D. Bruguera, Emilio L. Zapata
ASAP2
1996 High performance VLSI architecture for the trellis coded quantization
abstract
Trellis coded quantization (TCQ) is an efficient technique for encoding memoryless sources. Furthermore TCQ can be incorporated into a transform coding structure (such as the discrete cosine transform) for encoding monochrome and color images with fixed rate or entropy-constrained schemes. In all these cases an expanded codebook is partitioned into subsets used to label the branches of an appropriate graph (trellis). For a given data sequence, the Viterbi algorithm is then used to find the minimum mean square error path through the trellis. We present a generic architecture scheme that can be easily adapted to the different TCQ image compression methods. We also present a formal model that permits a regular and modular design solution that is optimal for a particular set of area and/or speed constraints.
Montserrat Bóo, Francisco Argüello, Javier D. Bruguera, Emilio L. Zapata
ICIP (2)2
1996 FFTs on Mesh Connected Computers
Francisco Argüello, Margarita Amor, Emilio L. Zapata
Parallel Comput.1
1995 A Parallel Architecture for the Self-Sorting FFT Algorithm
Francisco Argüello, Javier D. Bruguera, Emilio L. Zapata
J. Parallel Distributed Comput.1
1994 Parallel Architecture for Fast Transforms with Trigonometric Kernel
abstract
We present an unified parallel architecture for four of the most important fast orthogonal transforms with trigonometric kernel: Complex Valued Fourier (CFFT), Real Valued Fourier (RFFT), Hartley (FHT), and Cosine (FCT). Out of these, only the CFFT has a data flow coinciding with the one generated by the successive doubling method, which can be transformed on a constant geometry flow using perfect unshuffle or shuffle permutations. The other three require some type of hardware modification to guarantee the constant geometry of the successive doubling method. We have defined a generalized processing section (PS), based on a circular CORDIC rotator, for the four transforms. This PS section permits the evaluation of the CFFT and FCT transforms in n data recirculations and the RFFT and FHT transforms in n-1 data recirculations, with n being the number of stages of a transform of length N=r/sup n/. Also, the efficiency of the partitioned parallel architecture is optimum because there is no cycle loss in the systolic computation of all the butterflies for each of the four transforms.>
Francisco Argüello, Javier D. Bruguera, Ramón Doallo, Emilio L. Zapata
IEEE Trans. Parallel Distributed Syst.1
1992 A VLSI Constant Geometry Architecture for the Fast Hartley and Fourier Transforms
abstract
An application-specific architecture for the parallel calculation of the decimation in time and radix 2 fast Hartley (FHT) and Fourier (FFT) transforms is presented. A real sequence with N=2/sup n/ data items is considered as input. The system calculates the FHT and the FFT in n and n+1 stages. respectively. The modular and regular parallel architecture is based on a constant geometry algorithm using butterflies of four data items and the perfect unshuffle permutation. With this permutation, the mapping of the algorithm in VLSI technology is simplified and the communications among processors are minimized. Organization of the processor memory based on first-in, first-out (FIFO) queues facilitates a systolic data flow and permits the implementation in a direct way of the complex data movements and address sequences of the transforms. This is accomplished by means of simple multiplexing operations, using hardwired control. The total calculation time is (Nlog/sub 2/N)/4Q cycles for the FHT and N(1+log/sub 2/N)/4Q cycles for the FFT, where Q is the number of processors (Q= 2/sup q/, Q>
Emilio L. Zapata, Francisco Argüello
IEEE Trans. Parallel Distributed Syst.2
1991 Design of a constant geometry fast Hartley transformer
abstract
A semisystolic architecture is presented for the parallel calculation of the decimation in time and radix-2 fast Hartley transform (FHT) of a real sequence with N=2/sup n/ data items. The architecture is based on a constant geometry algorithm for computing the FHT which facilitates its mapping in VLSI technology and minimizes the communications among processors. The circuit proposed is characterized by its modular design and its interconnective regularity. It permits the computation of arbitrarily sized FHTs as a consequence of the partition of the data and the recirculation of partial results over the processing units in the successive stages of the transform. Each calculation stage requires N/4Q cycles where Q is the number of processors (Q=2/sup q/). The total calculation time is (Nlog/sub 2/N)/4Q cycles.>
Francisco Argüello, Ramón Doallo, Javier D. Bruguera, Emilio L. Zapata
ICASSP1
1990 Multidimensional fast Hartley transform onto SIMD hypercubes
Emilio L. Zapata, Francisco Argüello, Francisco F. Rivera, Javier D. Bruguera
Microprocessing and Microprogramming2
1990 Systolic architecture for the calculation of the correlation coefficients
Emilio L. Zapata, José Carlos Cabaleiro, Ramón Doallo, Francisco Argüello
Microprocessing and Microprogramming4