EDBT 2026 Demo / reviewers in the wild / expert
Roberto R. Osorio
dblp:60/6349
· DBLP profile ↗
21ranked-venue papers
14as first author
4since 2021 · last 2026
0000-0001-8768-2240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 10 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pipelined FPGA Implementation of a Differential Evolution Engine for Optimization of Scientific ModelsabstractCustom computing machines implemented on FPGAs have emerged as a powerful solution for tackling computationally intensive tasks, leveraging their capacity for deep pipelining and parallel memory access. Differential Evolution (DE), a robust optimization algorithm, combined with adaptive numerical integration methods, is widely used to optimize parameter values in diverse scientific models. These tasks involve extensive floating-point computations, making FPGAs an ideal platform for efficiently accelerating their execution. In this work, we present a flexible and scalable FPGA architecture optimized for DE. This architecture is tailored to solve complex, resource-intensive optimization problems and is easily customizable for various models and integration methods. To demonstrate its efficacy, we evaluate two case studies: The Hodgkin–Huxley model for neuron action potentials and the Circadian clock model of Arabidopsis thaliana . Our architecture integrates adaptive numerical methods with DE and achieves significant performance and energy efficiency gains over CPU and GPU implementations while maintaining versatility across applications. Our architecture’s modular design enables seamless adaptation across different scientific contexts, enabling further optimization of resource utilization and expansion of application domains. The results underline the potential of FPGAs as a superior platform for large-scale scientific computation, offering unmatched energy efficiency and computational throughput for highly demanding tasks. The code developed to carry out this work is publicly available at https://github.com/mdccUVa/de-fpga . Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2025 | Accelerating Scientific Model Optimization with a Pipelined FPGA-Based Differential Evolution EngineabstractCustom computing machines on FPGAs excel in solving computationally intensive tasks by leveraging deep pipelines and parallel memory access. This paper introduces a flexible FPGA-based architecture for Differential Evolution (DE), optimized for a variety of scientific models and numerical integration methods. The architecture's modular design allows seamless customization for diverse applications. Two case studies demonstrate the architecture's capabilities: the Hodgkin-Huxley model for neuron action potentials and the Circadian model of Arabidopsis thaliana. These implementations employ double-precision floating-point arithmetic and adaptive numerical integration techniques, addressing the challenges of complex, stiff differential equations. The proposed design outperforms CPUs and GPUs in computational speed and energy efficiency, achieving up to 3.8x faster processing and significant reductions in energy consumption. This work highlights the potential of FPGA platforms for accelerating complex scientific computations while providing insights for future optimizations in resource utilization and broader applicability. Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
FCCM | 2 |
| 2023 | Implementation of a motion estimation algorithm for Intel FPGAs using OpenCLabstractMotion Estimation is one of the main tasks behind any video encoder. It is a computationally costly task; therefore, it is usually delegated to specific or reconfigurable hardware, such as FPGAs. Over the years, multiple FPGA implementations have been developed, mainly using hardware description languages such as Verilog or VHDL. Since programming using hardware description languages is a complex task, it is desirable to use higher-level languages to develop FPGA applications.The aim of this work is to evaluate OpenCL, in terms of expressiveness, as a tool for developing this kind of FPGA applications. To do so, we present and evaluate a parallel implementation of the Block Matching Motion Estimation process using OpenCL for Intel FPGAs, usable and tested on an Intel Stratix 10 FPGA. The implementation efficiently processes Full HD frames completely inside the FPGA. In this work, we show the resource utilization when synthesizing the code on an Intel Stratix 10 FPGA, as well as a performance comparison with multiple CPU implementations with varying levels of optimization and vectorization capabilities. We also compare the proposed OpenCL implementation, in terms of resource utilization and performance, with estimations obtained from an equivalent VHDL implementation. Manuel de Castro, Roberto R. Osorio, David López Vilariño, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2023 | Floating Point Calculation of the Cube Function on FPGAsabstractSpecialized arithmetic units allow fast and efficient computation of lesser used mathematical functions. The overall impact of those units would be negligible in a general purpose processor, as added circuitry makes chips more complex despite most software would seldom make use of it. On the opposite side, custom computing machines are built for a specific task, and they can always benefit from specialized units if they are available. In this work, floating point architectures are proposed for computing the cube on Intel and Xilinx FPGAs. Those implementations reduce the cost and latency compared to using simple floating point multiplications and squarers. Roberto R. Osorio |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | A Microprogrammed Approach for Implementing StatechartsabstractStatechart diagrams allow specifying complex systems in which there may be several states active at the same time and a large number of events and transitions to evaluate. Statecharts have been found useful in the design and implementation of control systems in research facilities, such as particle accelerators. Automatic tools may convert statechart-based specifications into hardware descriptions. During the development of one of those tools, the convenience of implementing statecharts as microprogrammed control systems was considered. In this work, we propose a method for implementing generic microprogrammed architectures that support statecharts upgradable on the field. This approach is evaluated showing its advantages and disadvantages. Javier Cereijo García, Roberto R. Osorio |
DSD | 2 |
| 2016 | Pipelined FPGA implementation of numerical integration of the Hodgkin-Huxley modelabstractThe Hodgkin-Huxley model describes the initiation and propagation of action potential in neurons' axons. The model consists of a set of nonlinear differential equations that can be solved using numerical methods for a given choice of parameters. As the equations reflect physiological processes, the value of those parameters are subject to great variability. Therefore, numerical integration is often combined with differential evolution methods in order to find which set of parameters minimizes some fitness function. As modern FPGAs are large enough to implement complex functions using double-precision floating-point arithmetic, intensive scientific computations may be carried out showing competitive performance and cost. In this work, we present a pipelined architecture for performing the 4th order Runge-Kutta integration of the equations of the Hodgkin-Huxley model, introducing convenient implementations of complex mathematical functions. Roberto R. Osorio |
ASAP | 1 |
| 2016 | A fast algorithm for constructing nearly optimal prefix codesabstractSummary Huffman algorithm allows for constructing optimal prefix‐codes withO(n·logn) complexity. As the number of symbolsngrows, so does the complexity of building the code‐words. In this paper, a new algorithm and implementation are proposed that achieve nearly optimal coding without sorting the probabilities or building a tree of codes. The complexity is proportional to the maximum code length, making the algorithm especially attractive for large alphabets. The focus is put on achieving almost optimal coding with a fast implementation, suitable for real‐time compression of large volumes of data. A practical case example about checkpoint files compression is presented, providing encouraging results. Copyright © 2015 John Wiley & Sons, Ltd. Roberto R. Osorio, Patricia González |
Softw. Pract. Exp. | 1 |
| 2013 | Architecture and Implementation of a Data Compression System at Switch-Level in ATA-over-Ethernet Storage NetworksabstractIn this work, a new architecture for loss less data compression and decompression is integrated within an Ethernet switch using the NetFPGA open platform. The aim is compressing data packets in a block-based storage network. Data packets are compressed when written to the target disk and decompressed when read by the initiator. ATA-over-Ethernet (AoE) has been chosen as it is an efficient and relatively simple technology that does not rely on IP. The ultimate goal is achieving a better use of the available network bandwidth with the target and a possible reduction in power consumption. The use case of application-level check pointing in supercomputing is presented, for which compression ratios are given, and the efficiency of the proposed scheme is then discussed. Angela Souto Vieites, Roberto R. Osorio |
DSD | 2 |
| 2012 | Fast Construction of Nearly-Optimal Prefix Codes without Probability SortingabstractIn this abstract, an algorithm is proposed that achieves nearly-optimal coding without sorting the probabilities or building a tree of codes. The complexity is proportional to the maximum code length, making it especially attractive for large alphabets. Roberto R. Osorio, Patricia González |
DCC | 1 |
| 2009 | High Performance Image Processing on a Massively Parallel Processor ArrayabstractMulticore and many core processors are the new wave of computing, offering high performance by using large numbers of simple processors. In this paper, we describe the implementation of 2 applications into an Ambric massively parallel processor array from a hardware design point of view. An evaluation of performance and design effort is provided, showing that massive parallel processor arrays may challenges FPGAs in some applications. Roberto R. Osorio, Cesar Diaz-Resco, Javier D. Bruguera |
DSD | 1 |
| 2008 | An FPGA architecture for CABAC decoding in manycore systemsabstractArithmetic coding is an efficient entropy compression method that achieves results close to the entropy limit and it is used in modern standards such as JPEG-2000 and H.264. Arithmetic decoding (AD) in H.264 video coding standard is a sequential task that takes a significant part of computing time. In present and future multicore and manycore systems, AD becomes a bottleneck as it cannot be parallelized, limiting the concurrent execution of other tasks. In this paper, an FPGA-based accelerator is proposed to speed-up AD in H.264 and enable parallel decoding at macroblock and frame levels scaling up to tens or hundreds of cores. Roberto R. Osorio, Javier D. Bruguera |
ASAP | 1 |
| 2007 | Entropy Coding on a Programmable Processor Array for Multimedia SoCabstractEntropy encoding and decoding is a crucial part of any multimedia system that can be highly demanding in terms of computing power. Hardware implementation of typical compression and decompression algorithms is cumbersome, while conventional software implementations are slow due to bit-level operations, data dependencies and conditional branching. Several solutions have been proposed along the years, ranging from hardware accelerators for high-end systems to careful implementations in VLIW processors and instruction-set extensions, both hardwired and reconfigurable. Multimedia systems must often implement several encoders and decoders for different formats. Hence, a programmable solution is mandatory. However, programmable processors may be challenged by highly-complex algorithms. In this work, a highly efficient and low cost alternative is presented based on an array processor. The dataflow of several entropy coding algorithms has been studied, leading to the choice of an efficient programming model, processor layout and interconnection system. Results are presented for JPEG and H.264 image and video coding standards. Roberto R. Osorio, Javier D. Bruguera |
ASAP | 1 |
| 2006 | A Unified Architecture for H.264 Multiple Block-Size DCT with Fast and Low Cost QuantizationabstractAVC/H.264 is the new international standard for video coding jointly developed by ISO-MPEG and ITU-T, which offers a substantial compression gain when compared with H.263 and MPEG-4 simple profile. One of the main characteristics of H.264 is the introduction of a integer version of the discrete cosine transform initially applied to 4times4 pixels blocks, and later extended to 8times8 pixels for high quality video encoding. In this work, a unified architecture is proposed for parallel 8times8 integer DCT and iDCT, also able to process 4times4 DCT, iDCT and Hadamard transform. A very fast quantization/de-quantization scheme is presented based on prediction that allows parallel quantization with a single multiplier. This architecture also implements all-zero detection, eliminating coefficients with high cost as specified in the standard and anticipates entropy encoding. The proposed design has been synthesized in AMS 0.35mu technology and achieves a maximum speed of 67 MHz Javier D. Bruguera, Roberto R. Osorio |
DSD | 2 |
| 2006 | A Combined Memory Compression And Hierarchical Motion Estimation Architecture For Video Encoding In Embedded SystemsabstractIn this paper a new technique is presented that combines memory compression in video encoders with fast and efficient motion estimation (ME). This technique is mainly oriented to embedded systems, which demand simple and power aware algorithms. Video encoding needs increasing amounts of memory for storing reference pictures. Memory compression allows reducing the footprint of the application, lowering the total implementation cost. In this paper, we combine memory compression and hierarchical ME so that the overhead associated to implement both techniques is shared. Thus, a net gain in processing speed is obtained, while reducing costs and power consumption Roberto R. Osorio, Javier D. Bruguera |
DSD | 1 |
| 2006 | High-Throughput Architecture for H.264/AVC CABAC Compression SystemabstractNew image and video coding standards have pushed the limits of compression by introducing new techniques with high computational demands. The Advanced Video Coder (ITU-T H.264, AVC MPEG-4 Part 10) is the last international standard, which introduces new enhanced features that require new levels of performance. Among the new tools present in AVC, the context-based binary arithmetic coder (CABAC) offers significant compression advantage over baseline entropy coders. CABAC is meant to be used in AVC's Main and High Profiles, which target broadcast and video storage and distribution of standard and high-definition contents. In these applications, hardware acceleration is needed as the computational load of CABAC is high, challenging programmable processors. Moreover, rate-distortion optimization (RDO) increases CABAC's load by two orders of magnitude. In this paper, we present a fast and new architecture for arithmetic coding adapted to the characteristics of CABAC, including optimized use of memory and context managing and fast processing able to encode more than two symbols per cycle. A maximum processing speed of 185 MHz has been obtained for 0.35 mu, able to encode high quality video in real time. Some of the proposed optimization may also be applied to software implementations obtaining significant improvements Roberto R. Osorio, Javier D. Bruguera |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | A New Architecture for fast Arithmetic Coding in H.264 Advanced Video CoderabstractIn this work, a new architecture for binary arithmetic coding is presented in the context of the new AVC/H.264 standard for video coding. Among the new technologies included in AVC/H.264 a context adaptive binary arithmetic coder (CABAC) is used that outperforms the baseline entropy coder in a significant manner. In this work we justify the need for a new architecture that implements the unique characteristics of CABAC that are not found in other implementations of arithmetic coding. We show that a fast architecture is needed that combines short cycle time and application-aware scheduling in order to accomplish with the high computational demands. A number of optimizations are introduced that allow processing several symbols per cycle and reduce data binarization overhead. Implementation results are shown for a Virtex-II FPGA and the main conclusions are presented. Roberto R. Osorio, Javier D. Bruguera |
DSD | 1 |
| 2004 | Arithmetic Coding Architecture for H.264/AVC CABAC Compression SystemabstractIn this paper we propose an efficient implementation of CABAC's binary arithmetic coder and context management system. CABAC is the context adaptive binary arithmetic coder used in new H.264/AVC video standard. Arithmetic coding allows a significant enhancement in compression. However, implementation complexity is a drawback due to hardware cost and slowness. In this paper we show the need for a hardware implementation of arithmetic coding in current video compression systems. We propose a fast and efficient implementation of the encoding algorithm. We prove that memory accesses constitute a bottleneck and propose solutions that apply to the encoding algorithm and context management system. As a result, a fast architecture is presented, able to process one symbol per cycle. Roberto R. Osorio, Javier D. Bruguera |
DSD | 1 |
| 2004 | View-dependent, scalable texture streaming in 3-D QoS with MPEG-4 visual texture codingabstractMultimedia applications are characterized by high resource demands (computing power, memory, network bandwidth, and power consumption). Efficient implementations aim at reducing these resources to a minimum, which is of the utmost importance for small, low-cost terminals in low-bandwidth networks. Resource savings can also be obtained by content adaptation without impeding the quality of the decoded audio-visual media. In this context, the paper analyzes texture adaptation and streaming for three-dimensional applications, using MPEG-4's Visual Texture Coding tool, in conjunction with eXtensible Markup Language (XML)-based description techniques. Augmented features for content adaptation are supported, such as region selection, accompanied by resolution and SNR settings. As a result, quality is optimized for the terminal's computing capabilities and display resolution, taking the user's viewing conditions into account. Moreover, the instantaneous bandwidth utilization is highly reduced in streaming scenarios. Gauthier Lafruit, Eric Delfosse, Roberto R. Osorio, Wolfgang van Raemdonck, Vissarion Ferentinos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Bitstream Syntax Description Language for 3D MPEG-4 view-dependent texture streamingabstractIn modern multimedia applications, scalability is a key functionality that allows transmission and representation of content in a wide variety of networks and terminals. In order to obtain full advantage of scalability features, techniques for detailed description and transformation of multimedia contents are needed. In this paper, the Bitstream Syntax Description Language is used for describing the structure of an MPEG-4 wavelet-coded texture. An XML-based bitstream description transformation is then applied for selecting some texture regions at an appropriate quality, effectively scaling down the processing and bandwidth requirements for view-dependent texture transmission. Appropriately applying this technique to 3D streaming guarantees quality-of-service, i.e. it certifies the best quality at limited network/processing resources. Roberto R. Osorio, Sylvain Devillers, Eric Delfosse, Myriam Amielh, Gauthier Lafruit |
ICIP (3) | 1 |
| 1997 | New arithmetic coder/decoder architectures based on pipeliningabstractIn this paper we present new VLSI architectures for the arithmetic encoding and decoding of multilevel images. In these algorithms the speed is limited by their recursive natures and the arithmetic and memory access operations. They become specially critical in the case of decoding. In order to reduce the cycle length we propose working with two executions of the algorithm which alternate in the use of the pipelined hardware with a minimum increase in its cost. Roberto R. Osorio, Javier D. Bruguera |
ASAP | 1 |
| 1995 | Digit On-line Large Radix CORDIC RotatorabstractMany applications figure the evaluation of rotations at high speeds. However there is a trade-off between the chip area and the latency. In this paper we develop a digit on-line pipelined array architecture based on the radix-4 CORDIC algorithm in rotation mode. The radix-4 CORDIC algorithm halves the number of microrotations with respect the traditionally radix-2 algorithm with the drawback of a non-constant scale factor. Seeking a good compromise between silicon area and latency we have used digit on-line processing. This way the data inputs the processor in blocks of bits (digits) in MSD-first mode of processing. We have used redundant carry-save arithmetic to allow carry-free additions and on-line processing. The designed processor demonstrates to have a better performance than previous digit on-line architectures. Roberto R. Osorio, Elisardo Antelo, Javier D. Bruguera, Julio Villalba, Emilio L. Zapata |
ASAP | 1 |