EDBT 2026 Demo / reviewers in the wild / expert
Guillermo Botella Juan
dblp:61/7495 · also Guillermo Botella
· DBLP profile ↗
29ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0002-0848-2636ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Square Root Unit with Minimum Iterations for Posit ArithmeticabstractIn this paper, we introduce a novel implementation of a square root algorithm specifically tailored for posit arithmetic. Unlike traditional methods, the proposed approach capitalizes on the inherent flexibility of posits, which lack fixed-length fields, to optimize square root computations. By accurately estimating the minimum number of required fraction bits, our algorithm substantially reduces the recurrence iterations without sacrificing accuracy. Implemented across standard 16-bit, 32-bit, and 64-bit posit formats, our units showcase a significant latency reduction in different applications with only a marginal increase in resource utilization. Comparative analysis against previous pipelined designs underscores the area efficiency of our proposed solutions. This research significantly contributes to the advancement of posit-based arithmetic units, presenting promising opportunities for improving computational system efficiency. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ARITH | 3 |
| 2023 | A Suite of Division Algorithms for Posit ArithmeticabstractPosit ™ arithmetic is a promising alternative to IEEE 754 floating-point arithmetic due to its higher accuracy, larger dynamic range, and bitwise compatibility. While posit arithmetic has been well studied for basic arithmetic operations, division has received little attention. This paper proposes multiple divider designs for posit arithmetic based on digit recurrence and iterative approximation, and evaluates their performance. ASIC synthesis results show that the proposed designs significantly reduce hardware requirements for 32-bit division units compared to previous works by 1.14 × in area, 1.11× in power, and 1.04×in datapath delay. Moreover, the paper introduces an approximate logarithmic posit division that achieves an 8.8×reduction in area and 29×reduction in energy consumption with negligible degradation of the final results, making it suitable for error-tolerant applications. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ASAP | 3 |
| 2023 | PERCIVAL: Deploying Posits and Quire Arithmetic into the CVA6 RISC-V CoreabstractRepresenting and operating on real numbers in a microprocessor presents unique challenges not encountered with the set of integers. Working with real numbers introduces additional concepts such as precision, that is, the error made between the number with which we want to operate and the approximation that we can represent in a finite number of bits. Currently, the universally extended way of representing the set of real numbers is using floating-point numbers defined by the IEEE 754 standard. This format presents a series of difficulties, such as the different rounding schemes, reproducibility problems depending on the implementation, a multitude of ways to represent Not a Numbers (NaNs) or the existence of plus and minus zero. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
CF | 4 |
| 2023 | Generating Posit-Based Accelerators With High-Level SynthesisabstractRecently, the posit number system has demonstrated a higher accuracy over standard floating-point arithmetic for many scientific applications. However, when it comes to implementing accelerators for these applications, the tool support for this arithmetic format is still missing, especially during the step. In this paper, we incorporate the posit data type into the high-level synthesis (HLS) design process, so that we can generate the implementation directly from a given behavioral specification, but using posit numbers instead of the classical floating-point notations. Our evaluations show that, even if posit-based circuits require more area than their floating-point counterparts, they offer higher accuracy when using the same bitwidth. For example, using posit arithmetic can reduce computation errors by about two orders of magnitude when compared to using standard floating-point numbers. Our approach also includes an alternative to mitigate the high overheads of the posits and broadening the potential use of this format. We also propose a hybrid scheme that uses posit numbers only in the private local memory, while the accelerator operates in the classic floating-point notation. This solution is useful when the designers want to optimize local memories and data transfers, but still use legacy high-level synthesis (HLS) tools that only support traditional floating-point notations. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Christian Pilato |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | PERCIVAL: Open-Source Posit RISC-V Core With Quire CapabilityabstractPresents the front cover, title page, cover page, or splash screen of the proceedings record. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
ARITH | 4 |
| 2021 | Gender and STEAM as part of the MOOC STEAM4ALLabstractThis paper presents findings on participants of a massive open online course named "Educational Robotics for all" developed under Open edX platform. The document describes the organization and structure of the MOOC and some of its preliminary results. As an example, Module 2 about Gender and STEAM is presented and discussed. Carina Soledad González-González, Alicia García-Holgado, Pedro Plaza 0001, Manuel Castro 0001, Aruquia B. M. Peixoto, Julia Merino, Elio San Cristóbal, Antonio Menacho, Diana Urbano, Manuel Blázquez, Félix García Loro, Maria Teresa Restivo, Rebecca Strachan, Paloma Díaz 0001, Inmaculada Plaza, Cristina Fernández, Susan M. Lord, Diane T. Rover, Rosanna Yuen-Yan Chan, Melany M. Ciampi, Russ Meier 0001, Edmundo Tovar, Magdalena Salazar, Susan Zvacek, José A. Ruipérez-Valiente, Blanca Quintana, Sergio Martín 0001, Guillermo Botella Juan, África Lopez-Rey, Paulo Abreu |
EDUCON | 28 |
| 2021 | Energy-Efficient MAC Units for Fused Posit ArithmeticabstractPosit arithmetic is an alternative format to the standard IEEE 754 for floating-point numbers that claims to provide compelling advantages over floats, including higher accuracy, larger dynamic range, or bitwise compatibility across systems. The interest in the design of arithmetic units for this novel format has increased in the last few years. However, while multiple designs for posit adder and multiplier have been developed recently in the literature, fused units for posit arithmetic are still in the early stages of research. Moreover, due to the large size of accumulators needed in fused operations, the few fused posit units proposed so far still require many hardware resources. In order to contribute to the development of the posit number format, and facilitate its use in applications such as deep learning, this paper presents several designs of energy-efficient posit multiply- accumulate (MAC) units with support for standard quire format. Concretely, the proposed designs are capable of computing fused dot products of large vectors without accuracy drop, while consuming less energy than previous implementations. Experiments show that, compared to previous implementations, the proposed designs consume up to 75.49%, 88.45% and 83.43% less energy and are 73.18%, 87.36% and 83.00% faster for 8, 16 and 32 bitwidths, with an additional area of only 4.97%, 7.44% and 4.24%, respectively. Raul Murillo 0001, David Mallasén, Alberto A. Del Barrio, Guillermo Botella Juan |
ICCD | 4 |
| 2021 | First experiences of teaching quantum computing
Ginés Carrascal, Alberto A. Del Barrio, Guillermo Botella Juan |
J. Supercomput. | 3 |
| 2020 | Customized Posit Adders and Multipliers using the FloPoCo Core GeneratorabstractThe posit number system, which is proposed as a replacement of IEEE floating-point numbers, is in the spotlight of Arithmetic research due to the recent breakthroughs. This format claims to provide more accurate results with the same bitwidth than standard floating point, but the run-time variability during the detection of the posit fields involves a hardware design challenge. In this work, we propose parameterized designs for multiple posit functional units, including addition and multiplication, and integrate them as templates of the FloPoCo framework. The integration of the proposed algorithms within FloPoCo can provide synthesizable VHDL code for posit arithmetic of any possible configuration 〈n, es〉. Experiments show an improvement in terms of area and energy with respect to state-of-the-art works up to 35.9% and 30.8%, respectively. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ISCAS | 3 |
| 2020 | HEVC optimization based on human perception for real-time environments
David Guillermo Fernández, Guillermo Botella Juan, Alberto A. Del Barrio, Carlos García 0001, Manuel Prieto 0001, Christos Grecos |
Multim. Tools Appl. | 2 |
| 2019 | Portability Study of an OpenCL Algorithm for Automatic Target Detection in Hyperspectral ImagesabstractIn the last decades, the problem of target detection has received considerable attention in remote sensing applications. When this problem is tackled using hyperspectral images with hundreds of bands, the use of high-performance computing (HPC) is essential. One of the most popular algorithms in the hyperspectral image analysis community for this purpose is the automatic target detection and classification algorithm (ATDCA). Previous research has already investigated the mapping of ATDCA on HPC platforms such as multicore processors, graphics processing units (GPUs), and field-programmable gate arrays (FPGAs), showing impressive speedup factors (after careful fine-tuning) that allow for its exploitation in time-critical scenarios. However, the lack of standardization resulted in most implementations being too specific to a given architecture, eliminating (or at least making extremely difficult) code reusability across different platforms. In order to address this issue, we present a portability study of an implementation of ATDCA developed using the open computing language (OpenCL). We focus on cross-platform parameters such as performance, energy consumption, and code design complexity, as compared to previously developed (hand-tuned) implementations. Our portability study analyzes different strategies to expose data parallelism as well as enable the efficient exploitation of complex memory hierarchies in heterogeneous devices. We also conduct an assessment of energy consumption and discuss metrics to analyze the quality of our code. The conducted experiments-using synthetic and real hyperspectral data sets collected by the Hyperspectral Digital Imagery Collection Experiment (HYDICE) and NASA's Airborne Visible Infra-Red Imaging Spectrometer (AVIRIS)-demonstrate, for the first time in the literature, that portability across different HPC platforms can be achieved for real-time target detection in hyperspectral missions. Sergio Bernabé, Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | A Fast Image Dehazing Algorithm Using Morphological ReconstructionabstractOutdoor images are used in a vast number of applications, such as surveillance, remote sensing, and autonomous navigation. The greatest issue with these types of images is the effect of environmental pollution: haze, smog, and fog originating from suspended particles in the air, such as dust, carbon and water drops, which cause degradation to the image. The elimination of this type of degradation is essential for the input of computer vision systems. Most of the state-of-the-art research in dehazing algorithms is focused on improving the estimation of transmission maps, which are also known as depth maps. The transmission maps are relevant because they have a direct relation to the quality of the image restoration. In this paper, a novel restoration algorithm is proposed using a single image to reduce the environmental pollution effects, and it is based on the dark channel prior and the use of morphological reconstruction for the fast computing of transmission maps. The obtained experimental results are evaluated and compared qualitatively and quantitatively with other dehazing algorithms using the metrics of the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) index; based on these metrics, it is found that the proposed algorithm has improved performance compared to recently introduced approaches. Sebastián Salazar-Colores, Eduardo Cabal-Yepez, Juan Manuel Ramos-Arreguín, Guillermo Botella Juan, Luis Manuel Ledesma-Carrillo, Sergio E. Ledesma-Orozco |
IEEE Trans. Image Process. | 4 |
| 2018 | Intra-Steganography: Hiding Data in High-Resolution VideosabstractSteganography is the art of hiding information within a file like an image or a video. The embedded message can then be used either to transmit some secret information or to protect the content of the file. Due to the increase of the resolutions to provide higher quality in videos, it is critical to comply with the latest video standard, namely: the High Efficiency Video Coding (HEVC), which allows reducing the size of the file to be transmitted. Thus, in this paper we propose an HEVC-compliant method to hide and retrieve information in high-resolution videos. The procedure is based on modifying the luminance of certain blocks. Nevertheless, this must be carefully done, as the HEVC standard is a powerful attack in itself, since it compresses 50% the size of the video on average and the embedded information may disappear. In this paper there is a study evaluating a proper spot to embed the information. Results show that it is possible to retrieve all the information while maintaining the quality of the video after embedding the message. Several tests have been run with the reference software HM-16.2 as well as the real-time encoder x265, and the SSIM and PSNR values are coherent with these theses. David Rodríguez 0002, Alberto A. Del Barrio, Guillermo Botella Juan, David Cuesta |
DS-RT | 3 |
| 2018 | Fast and effective CU size decision based on spatial and temporal homogeneity detection
David Guillermo Fernández, Alberto A. Del Barrio, Guillermo Botella Juan, Carlos García 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Embedded Grammars for Grammatical Evolution on GPGPU
J. Ignacio Hidalgo, Carlos Cervigón, José Manuel Velasco, José Manuel Colmenar, Carlos García 0001, Guillermo Botella Juan |
EvoApplications (1) | 6 |
| 2017 | First Experiences Accelerating Smith-Waterman on Intel's Knights Landing Processor
Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
ICA3PP | 3 |
| 2017 | Performance-Power Evaluation of an OpenCL Implementation of the Simplex Growing Algorithm for Hyperspectral UnmixingabstractOver the last few years, several new strategies for spectral unmixing of remotely sensed hyperspectral data have been proposed. Many of them have been developed to solve the most time-consuming and relevant step: endmember extraction. However, unmixing algorithms can be computationally very expensive in terms of processing time and energy consumption, a fact that compromises their use in applications under real-time and energy/power constraints. In this letter, we present a new parallel simplex growing algorithm (SGA) for hyperspectral data which exploits the memory hierarchy with operations in single-precision floating point. Those optimizations accelerate the most time-consuming parts of this method using the open computing language (OpenCL) standard. We have evaluated the performance versus energy consumption using the same open standard for parallel programming over a diverse set of heterogeneous platforms. Experiments have been conducted using real hyperspectral images collected by NASA's Airborne Visible Infrared Imaging Spectrometer and a collection of 24 synthetic hyperspectral images simulated with different sizes and number of endmembers (10-30). Considering the power consumption and OpenCL across all the proposed devices, the analysis presented indicates that the SGA can now be executed in computationally efficient fashion, which was not possible before introducing the parallel implementation described in this letter. Sergio Bernabé, Guillermo Botella Juan, Jose M. R. Navarro, Carlos Orueta, Francisco D. Igual, Manuel Prieto 0001, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Parallel implementation of the simplex growing algorithm for hyperspectral unmixing using OpenCLabstractMany algorithms for spectral unmixing have been proposed in the last years applied on hyperspectral imaging. This process is composed by three stages where the extraction of endmembers is the most consuming step. However, endmember extraction algorithms (EEAs) can be computationally very expensive and its acceleration on parallel architectures is still an interesting and open problem. In this paper, we present a parallel implementation of the simplex growing algorithm for hyperspectral unmixing called P-SGA on different platforms using the OpenCL framework. The proposed implementation exploits the memory hierarchy to accelerate the parts of this method which are more time-consuming. The proposed algorithm is evaluated in terms of both accuracy and computational performance through Monte Carlo simulations using the following architectures: multi-core Xeon CPU, NVidia GeForce GTX 980 GPU and Intel Xeon Phi accelerator. Experiments are conducted using real hyperspectral data set revealing considerable acceleration factors, which satisfies the real-time constraints given by the data acquisition rate. Sergio Bernabé, Guillermo Botella Juan, Jose M. R. Navarro, Carlos Orueta, Manuel Prieto 0001, Antonio Plaza |
IGARSS | 2 |
| 2015 | Early Experiences with OpenCL on FPGAs: Convolution Case StudyabstractMany proprietary standards and tools have been designed in order to cover a closed set of architectures, and OpenCL has become a free standard for parallel programming on heterogeneous systems, which include custom devices, CPUs, GPUs, FPGAs. This work evaluates the use of the well-known convolution operator in signal processing disciplines focused on FPGA evaluation under different optimizations with respect to thread and memory level exploitation. Carlos Rodriguez-Donate, Guillermo Botella Juan, Carlos García 0001, Eduardo Cabal-Yepez, Manuel Prieto 0001 |
FCCM | 2 |
| 2015 | An energy-aware performance analysis of SWIMM: Smith-Waterman implementation on Intel's Multicore and Manycore architecturesabstractSummary Alignment is essential in many areas such as biological, chemical and criminal forensics. The well‐known Smith–Waterman (SW) algorithm is able to retrieve the optimal local alignment with quadratic time and space complexity. There are several implementations that take advantage of computing parallelization, such as manycores, FPGAs or GPUs, in order to reduce the alignment effort. In this research, we adapt, develop and tune the SW algorithm named SWIMM on a heterogeneous platform based on Intel's Xeon and Xeon Phi coprocessor. SWIMM is a free tool available in a public git repository https://github.com/enzorucci/SWIMM . We efficiently exploit data and thread‐level parallelism, reaching up to 380 GCUPS on heterogeneous architecture, 350 GCUPS for the isolated Xeon and 50 GCUPS on Xeon Phi. Despite the heterogeneous implementation obtaining the best performance, it is also the most energy‐demanding. In fact, we also present a trade‐off analysis between performance and power consumption. The greenest configuration is based on an isolated multicore system that exploits AVX2 instruction set architecture reaching 1.5 GCUPS/Watts. Copyright © 2015 John Wiley & Sons, Ltd. Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | Smith-Waterman algorithm on heterogeneous systems: A case studyabstractThe well-known Smith-Waterman (SW) algorithm is a high-sensitivity method for local alignments. However, SW is expensive in terms of both execution time and memory usage, which makes it impractical in many applications. Some heuristics are possible but at the expense of losing sensitivity. Fortunately, previous research have shown that new computing platforms such as GPUs and FPGAs are able to accelerate SW and achieve impressive speedups. In this paper we have explored SW acceleration on a heterogeneous platform equipped with an Intel Xeon Phi coprocessor. Our evaluation, using the well-known Swiss-Prot database as a benchmark, has shown that a hybrid CPU-Phi heterogeneous system is able to achieve competitive performance (62.6 GCUPS), even with moderate low-level optimisations. Enzo Rucci, Armando De Giusti, Marcelo R. Naiouf, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001 |
CLUSTER | 4 |
| 2013 | Non-negative matrix factorization on low-power architectures: a comparative studyabstractPower consumption is emerging as one of the main concerns in the High Performance Computing (HPC) field. Many bioinformatics applications require HPC techniques and parallel architectures to meet performance requirements, but at the same time they can be severely limited by energy consumption restrictions. In this paper, we perform an empirical study of an optimized implementation of the Nonnegative Matrix Factorization (NMF), that is widely used in many fields of bioinformatics. We target different types of architectures, including general-purpose, low-power embedded processors and specific-purpose architectures like graphics processors and digital signal processors. From our study, we gain insights in both performance and energy consumption for each one of them under given experimental conditions, and conclude that the most appropriate architecture is usually a trade-off between performance and power consumption for a given experiment and dataset. Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Francisco Tirado |
EuroMPI | 3 |
| 2013 | GPU-based acceleration of bio-inspired motion estimation modelabstractSUMMARY In this paper, we describe the specific and efficient implementation of a gradient‐based optical flow model. This scheme was particularized using a validated neuromorphic motion estimation system for the robust extraction of image velocity. This model contains many characteristics that enhanced the capability when compared with other optical flow gradient family algorithms. Our implementation was performed using specific graphic processing units designed in an ad hoc framework for this model, which could be reused in several low‐level machine‐vision approaches. Observed performance results indicate that these accelerators be highly recommended. Furthermore, the throughput obtained in comparison with a general CPU was analyzed for the accurateness of a system built with regard to other optical flow systems. Additionally, several visual examples, commonly used for testing motion estimation sequences, were shown to reveal implementation behavior features. Copyright © 2012 John Wiley & Sons, Ltd. Fermin Ayuso, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001, Francisco Tirado |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Stochastic stability analysis of competitive neural networks with different time-scales
Anke Meyer-Bäse, Guillermo Botella Juan, Liliana Rybarska-Rusinek |
Neurocomputing | 2 |
| 2012 | Dyna-H: A heuristic planning reinforcement learning algorithm applied to role-playing game strategy decision systems
Matilde Santos Peñas, José Antonio Martín H., Victoria López, Guillermo Botella Juan |
Knowl. Based Syst. | 4 |
| 2011 | FPGA-Based Acceleration of Block Matching Motion Estimation TechniquesabstractThis paper focuses on the hardware acceleration of Block Matching motion estimation techniques (Search reduction family) suitable for the standard H.264/AVC MPEG-4 part 10 video compression. Many representative motion estimation search algorithms are explained here. As hardware, the well known Altera DE2 platform with a Cyclone II EP2C35F672C6 is used with a soft core NIOS II processor. C2H compiler which permits us to speed up the system at least two magnitude order is used to accelerate our source code. The paper shows the results in terms of performance and resources needed. This is the starting point to accelerate motion estimation algorithms using many strategies considering an ad-hoc motion estimation processor. Diego González 0002, Guillermo Botella Juan, Soumak Mokheerje, Uwe Meyer-Bäse |
FPL | 2 |
| 2010 | Robust Bioinspired Architecture for Optical-Flow ComputationabstractMotion estimation from image sequences, called optical flow, has been deeply analyzed by the scientific community. Despite the number of different models and algorithms, none of them covers all problems associated with real-world processing. This paper presents a novel customizable architecture of a neuromorphic robust optical flow (multichannel gradient model) based on reconfigurable hardware with the properties of the cortical motion pathway, thus obtaining a useful framework for building future complex bioinspired real-time systems with high computational complexity. The presented architecture is customizable and adaptable, while emulating several neuromorphic properties, such as the use of several information channels of small bit width, which is the nature of the brain. This paper includes the resource usage and performance data, as well as a comparison with other systems. This hardware platform has many application fields in difficult environments due to its bioinspired nature and robustness properties, and it can be used as starting point in more complex systems. Guillermo Botella Juan, Antonio García 0001, M. Rodriguez-Alvarez, Eduardo Ros Vidal, Uwe Meyer-Bäse, María C. Molina |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Enhanced gradient-based motion vector coprocessorabstractThe present work describes a reliable improved gradient optical flow estimation system using FPGAs. This structure is based on space-temporal processing and the use of steerable filters. This model can be enhanced using psychophysical and bioinspired properties according to biological vision in order to mimic the singularity and the performance of mammalians. Experimental results and the resources used to analyze the associate customizability of the system are discussed. Guillermo Botella Juan, Antonio García 0001, Uwe Meyer-Bäse, Manuel Rodríguez 0002, María C. Molina, Luís Parrilla Roure |
FPL | 1 |
| 2008 | Exploiting Internal Operation Patterns during the High-Level Synthesis of Time-Constrained CircuitsabstractConventional high-level synthesis algorithms treat specification operations as atomic elements that are executed in one or several consecutive cycles and over one functional unit. However, in most specifications there exist different operations, in function of their type, representation, and width that handled at different decomposition levels may produce better designs. In this way, most arithmetic operations can be decomposed into smaller operations applying several arithmetical properties. Different decompositions can be performed, in function of the pursued objective: performance improvement, area reduction, or power consumption reduction. In this paper we propose a pattern-based design methodology able to treat every operation at its most appropriate decomposition level. It produces reduced datapaths while meeting the specified time constraints. In comparison to conventional algorithms the amount of area saved averages 40%. Pedro Garcia-Repetto, María C. Molina, Rafael Ruiz-Sautua, Guillermo Botella Juan |
DSD | 4 |