Manuel de Castro

dblp:93/2300 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-3080-5136ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Pipelined FPGA Implementation of a Differential Evolution Engine for Optimization of Scientific Models
abstract
Custom computing machines implemented on FPGAs have emerged as a powerful solution for tackling computationally intensive tasks, leveraging their capacity for deep pipelining and parallel memory access. Differential Evolution (DE), a robust optimization algorithm, combined with adaptive numerical integration methods, is widely used to optimize parameter values in diverse scientific models. These tasks involve extensive floating-point computations, making FPGAs an ideal platform for efficiently accelerating their execution. In this work, we present a flexible and scalable FPGA architecture optimized for DE. This architecture is tailored to solve complex, resource-intensive optimization problems and is easily customizable for various models and integration methods. To demonstrate its efficacy, we evaluate two case studies: The Hodgkin–Huxley model for neuron action potentials and the Circadian clock model of Arabidopsis thaliana . Our architecture integrates adaptive numerical methods with DE and achieves significant performance and energy efficiency gains over CPU and GPU implementations while maintaining versatility across applications. Our architecture’s modular design enables seamless adaptation across different scientific contexts, enabling further optimization of resource utilization and expansion of application domains. The results underline the potential of FPGAs as a superior platform for large-scale scientific computation, offering unmatched energy efficiency and computational throughput for highly demanding tasks. The code developed to carry out this work is publicly available at https://github.com/mdccUVa/de-fpga .
Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris
ACM Trans. Reconfigurable Technol. Syst.1
2025 Accelerating Scientific Model Optimization with a Pipelined FPGA-Based Differential Evolution Engine
abstract
Custom computing machines on FPGAs excel in solving computationally intensive tasks by leveraging deep pipelines and parallel memory access. This paper introduces a flexible FPGA-based architecture for Differential Evolution (DE), optimized for a variety of scientific models and numerical integration methods. The architecture's modular design allows seamless customization for diverse applications. Two case studies demonstrate the architecture's capabilities: the Hodgkin-Huxley model for neuron action potentials and the Circadian model of Arabidopsis thaliana. These implementations employ double-precision floating-point arithmetic and adaptive numerical integration techniques, addressing the challenges of complex, stiff differential equations. The proposed design outperforms CPUs and GPUs in computational speed and energy efficiency, achieving up to 3.8x faster processing and significant reductions in energy consumption. This work highlights the potential of FPGA platforms for accelerating complex scientific computations while providing insights for future optimizations in resource utilization and broader applicability.
Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris
FCCM1
2023 Implementation of a motion estimation algorithm for Intel FPGAs using OpenCL
abstract
Motion Estimation is one of the main tasks behind any video encoder. It is a computationally costly task; therefore, it is usually delegated to specific or reconfigurable hardware, such as FPGAs. Over the years, multiple FPGA implementations have been developed, mainly using hardware description languages such as Verilog or VHDL. Since programming using hardware description languages is a complex task, it is desirable to use higher-level languages to develop FPGA applications.The aim of this work is to evaluate OpenCL, in terms of expressiveness, as a tool for developing this kind of FPGA applications. To do so, we present and evaluate a parallel implementation of the Block Matching Motion Estimation process using OpenCL for Intel FPGAs, usable and tested on an Intel Stratix 10 FPGA. The implementation efficiently processes Full HD frames completely inside the FPGA. In this work, we show the resource utilization when synthesizing the code on an Intel Stratix 10 FPGA, as well as a performance comparison with multiple CPU implementations with varying levels of optimization and vectorization capabilities. We also compare the proposed OpenCL implementation, in terms of resource utilization and performance, with estimations obtained from an equivalent VHDL implementation.
Manuel de Castro, Roberto R. Osorio, David López Vilariño, Arturo González-Escribano, Diego R. Llanos Ferraris
J. Supercomput.1
2023 EPSILOD: efficient parallel skeleton for generic iterative stencil computations in distributed GPUs
abstract
Abstract Iterative stencil computations are widely used in numerical simulations. They present a high degree of parallelism, high locality and mostly-coalesced memory access patterns. Therefore, GPUs are good candidates to speed up their computation. However, the development of stencil programs that can work with huge grids in distributed systems with multiple GPUs is not straightforward, since it requires solving problems related to the partition of the grid across nodes and devices, and the synchronization and data movement across remote GPUs. In this work, we present EPSILOD, a high-productivity parallel programming skeleton for iterative stencil computations on distributed multi-GPUs, of the same or different vendors that supports any type of n-dimensional geometric stencils of any order. It uses an abstract specification of the stencil pattern (neighbors and weights) to internally derive the data partition, synchronizations and communications. Computation is split to better overlap with communications. This paper describes the underlying architecture of EPSILOD, its main components, and presents an experimental evaluation to show the benefits of our approach, including a comparison with another state-of-the-art solution. The experimental results show that EPSILOD is faster and shows good strong and weak scalability for platforms with both homogeneous and heterogeneous types of GPU.
Manuel de Castro, Inmaculada Santamaria-Valenzuela, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris
J. Supercomput.1