VLDB 2026 Research / reviewers in the wild / expert
Andrés Otero
dblp:82/7913 · also J. Andrés Otero
· DBLP profile ↗
23ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-4995-7009ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accuracy-Performance-Resources Trade-Offs in RISC-V Microarchitectures for Genetic ProgrammingabstractAmong machine learning techniques, Genetic Programming (GP) flexibly adapts algorithm topology and complexity to the target problem. This adaptability makes GP models computationally efficient at inference, requiring fewer resources than many alternatives. This property aligns well with resource-constrained embedded systems with diverse performance and energy requirements. Paul Allaire, Mickaël Dardaillon, Thibaut Marty, Alfonso Rodríguez 0002, Andrés Otero, Karol Desnos |
CF | 5 |
| 2026 | From Cloud-Heavy to Edge-Ready: Self-supervised Transfer-efficient Emotion RecognitionabstractDeploying AI-based emotion recognition at the edge enables real-world applications but is constrained by data scarcity, heavy models, hardware limits, and privacy issues. To overcome these, we propose CHEER (Cloud-HEavy to Edge-Ready), a self-supervised, transfer-efficient framework where the cloud pre-trains lightweight graph-based encoders using unlabeled data, stores them as frozen models, and deploys only the needed encoder. New users are locally matched via centroids of clusters, and a small on-device classifier is trained with minimal labeled data, reducing computation, memory, and energy use while preserving privacy. Experimental results in the WEMAC and WESAD datasets show an accuracy of 78.19% and 80.08% at the edge on a NVIDIA Jetson Orin Nano. Moreover, CHEER achieves more than 60% reduction in model size, and lowers both peak RAM usage and energy consumption by more than 50% compared to the state-of-the-art. Junjiao Sun, José Miranda 0001, Jorge Portilla, Andrés Otero |
DATE | 4 |
| 2026 | LsmGrid: A Data-Aware Neuroevolution Framework for Designing Minimal Liquid State MachinesabstractLiquid State Machines (LSM) are spiking recurrent neural networks inspired by the dynamics of the human brain. By encoding information as discrete spikes, they enable low-power and real-time data processing at the edge. Despite the benefits of this computational model, the design of efficient LSM still lacks a widely accepted methodological foundation. Existing approaches often rely on arbitrary parameter tuning and large reservoir sizes, which are computationally costly and demand prior domain-specific knowledge. In this work, we introduce a systematic methodology for simplifying LSM design while substantially reducing the number of required neurons and synapses. This significantly reduces power consumption, bringing this class of models closer to their original motivation of ultra-low-power processing directly at the sensor level. A hybrid metaheuristic combining hill climbing and genetic programming (HC-GP) optimizes the LSM design, guiding the network to efficiently capture input patterns while avoiding random connectivity inefficiencies. Experimental results on the N-TIDIGITS and FSDD benchmarks demonstrate the effectiveness of the approach, achieving 82.3% accuracy on N-TIDIGITS, significantly surpassing existing models of comparable size, and 92.7% on FSDD, achieving state-of-the-art accuracy with only 66 neurons (6.6% of the neurons reported in previous state-of-the-art networks). Andrés Ignacio Romo, Andrés Otero, José Manuel Lanza-Gutiérrez |
GECCO | 2 |
| 2025 | Solving the Cold-Start Problem for the Edge: Clustering and Adaptive Deep Learning for Emotion DetectionabstractDesigning AI-based applications personalized to each user's behavior presents significant challenges due to the cold start problem and the impracticality of extensive individual data labeling. These challenges are further compounded when deploying such applications at the edge, where limited computing resources constrain the design space. This paper introduces a novel approach to AI-driven personalized solutions in biosensing applications by combining deep learning with clustering-based separation techniques. The proposed Clustering and Learning for Emotion Adaptive Recognition (CLEAR) methodology strikes a balance between population-wide models and fully personalized systems by leveraging data-driven clustering. CLEAR demonstrates its effectiveness in emotion recognition tasks, and its integration with fine-tuning enables efficient deployment on edge devices, ensuring data privacy and real-time detection when new users are introduced to the system. We conducted experiments for model personalization on two edge computing platforms: the Coral Edge TPU Dev Board and the Raspberry Pi with an Intel Movidius Neural Compute Stick 2. The results show that initial cluster assignment for new users can be achieved without labeled data, directly addressing the cold-start problem. Compared to baseline validation without clustering, this proposal improves accuracy metric from 75% to 81.9%. Furthermore, fine-tuning with minimal labeled data significantly improves accuracy, achieving up to 86.34% for the fear detection task in the WEMAC dataset while remaining suitable for deployment on resource-constrained edge devices. Junjiao Sun, Laura Gutiérrez-Martín, Celia López-Ongil, José Miranda 0001, Jorge Portilla, Andrés Otero |
DATE | 6 |
| 2025 | Hardware-aware Decision Tree Ensembles for Radar Data Processing on the Edge with Online LearningabstractEdge machine learning solutions supporting realtime decision making require inference engines that combine high accuracy, low latency, and energy efficiency. While decision tree ensembles such as those generated with XGBoost and LightGBM remain widely used due to their interpretability and predictive performance, their increasing ensemble sizes pose challenges for hardware deployment. This paper introduces a novel AI-to-Hardware co-design framework that jointly optimizes the model structure and the hardware architecture to enable efficient decision tree inference at the edge. Instead of relying on conventional gradient-boosted training, the proposed approach employs a custom evolutionary algorithm (EA) to evolve compact yet accurate decision tree ensembles, significantly reducing model complexity. A dedicated decision tree inference accelerator is developed and integrated into the ESP (Embedded Scalable Platform) RISC-V System-on-Chip (SoC) platform, supporting runtime model reconfiguration, hardware-in-the-loop training, and online learning capabilities. Experimental validation on radar and electronic warfare classification tasks demonstrates that the evolved models achieve near-perfect accuracy $(99-100 \%)$ with an order of magnitude fewer trees. Overall, the hardware accelerator yields up to a $9 \times$ reduction in inference latency compared to a CPU baseline. The results highlight the effectiveness of co-optimizing learning algorithms and hardware architectures to meet edge intelligence’s stringent performance and efficiency requirements and the potential of open hardware technology to support real industrial applications. Rodrigo Olmos, Pedro Lobo, Andrés Otero, Sergio Hernández, Eduardo Casanueva |
DSD | 3 |
| 2025 | Leveraging Incremental Machine Learning for Reconfigurable Systems Modeling under Dynamic WorkloadsabstractDynamic workload orchestration is one of the main concerns when working with heterogeneous computing infrastructures in the edge-cloud continuum. In this context, FPGA-based computing nodes can take advantage of their improved flexibility, performance, and energy efficiency provided that they use proper resource management strategies. In this regard, many state-of-the-art systems rely on proactive power management techniques and task scheduling decisions, which in turn require deep knowledge about the applications to be accelerated and the actual response of the target reconfigurable fabrics when executing them. While acquiring this knowledge at design time was more or less feasible in the past, with applications mostly being static task graphs that did not change at run time, the highly dynamic nature of current workloads in the edge-cloud continuum, where tasks can be deployed on any node and at any time, has removed this possibility. As a result, being able to derive such information at run time to make informed decisions has become a must. This article presents an infrastructure to build incremental ML models that can be used to obtain run-time power consumption and performance estimations in FPGA-based reconfigurable multi-accelerator systems operating under dynamic workloads. The proposed infrastructure features a novel stop-and-restart resource-aware mechanism to monitor and control the model training and evaluation stages during normal system operation, enabling low-overhead updates in the models to account for either unexpected acceleration requests (i.e., tasks not considered previously by the models) or model drift (e.g., fabric degradation). Experimental results show that the proposed approach induces a maximal additional error of 3.66% compared to a continuous training alternative. Furthermore, the proposed approach incurs only a 4.49% execution time overhead, compared to the 20.91% overhead induced by the continuous training alternative. The proposed modeling strategy enables innovative scheduling approaches in reconfigurable systems. This is exemplified by the conflict-aware scheduler introduced in this work, which achieves up to a 1.35 times speedup in executing the experimental workload. Additionally, the proposed approach demonstrates superior adaptability compared to other methods in the literature, particularly in response to significant changes in workload and to mitigate the effects of model overfitting. The portability of the proposed modeling methodology and monitoring infrastructure is also shown through their application to both Zynq-7000 and Zynq UltraScale+ devices. Juan Encinas, Alfonso Rodríguez 0002, Andrés Otero |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Negative emotion recognition based on physiological signals using a CNN-LSTM modelabstractNegative emotions can lead to a variety of physiological and psychological problems. Identifying and interpreting negative emotions can help people and specialists to deal with their effects on the human body. This paper introduces a deep learning-based method for recognizing negative emotions utilizing physiological signals. It is based on extracting a set of 123 features from raw signals and organizing them into 2D feature maps. A hybrid CNN-LSTM model is then employed to classify these maps, simultaneously learning and integrating integral and sequential features to output emotional feedback. Feature fusion is incorporated during training to improve the method’s accuracy and reduce the data complexity. The approach has been validated using WEMAC and WESAD datasets, achieving F1-scores of 85.15% and 91.44% and accuracies of 85.03% and 90.98%, respectively. Furthermore, the feasibility of this method for real-world applications is demonstrated through deployment on the Coral Edge TPU, an embedded device, indicating its potential for run-time decision-making on the computing edge. Junjiao Sun, Jorge Portilla, Andrés Otero |
BIBM | 3 |
| 2024 | A Deep Learning Approach for Fear Recognition on the Edge Based on Two-Dimensional Feature MapsabstractApplying affective computing techniques to recognize fear and combining them with portable signal monitors makes it possible to create real-time detection systems that could act as bodyguards when users are in danger. With this aim, this paper presents a fear recognition method based on physiological signals obtained from wearable devices. The procedure involves creating two-dimensional feature maps from the raw signals, using data augmentation and feature selection algorithms, followed by deep learning-based classification models, taking inspiration from those used in image processing. This proposal has been validated with two different datasets, achieving, in WEMAC, WESAD 3-classes, and WESAD 2-classes, F1-score results of 78.13%, 88.07%, and 99.60%, respectively, and 79.90%, 89.12%, and 99.60% in accuracy. Furthermore, the paper demonstrates the feasibility of implementing the proposed method on the Coral Edge TPU device, prepared to make inferences on the edge. Junjiao Sun, Jorge Portilla, Andrés Otero |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | A Framework for Fast Prototyping of Photo-realistic Environments with Multiple PedestriansabstractRobotic applications involving people often require advanced perception systems to better understand complex real-world scenarios. To address this challenge, photo-realistic and physics simulators are gaining popularity as a means of generating accurate data labeling and designing scenarios for evaluating generalization capabilities, e.g., lighting changes, camera movements or different weather conditions. We develop a photo-realistic framework built on Unreal Engine and AirSim to generate easily scenarios with pedestrians and mobile robots. The framework is capable to generate random and customized trajectories for each person and provides up to 50 ready-to-use people models along with an API for their metadata retrieval. We demonstrate the usefulness of the proposed framework with a use case of multi-target tracking, a popular problem in real pedestrian scenarios. The notable feature variability in the obtained perception data is presented and evaluated. Sara Casao, Andrés Otero, Álvaro Serra-Gómez, Ana Cristina Murillo, Javier Alonso-Mora, Eduardo Montijano |
ICRA | 2 |
| 2022 | Exploiting Hardware-Based Data-Parallel and Multithreading Models for Smart Edge Computing in Reconfigurable FPGAsabstractCurrent edge computing systems are deployed in highly complex application scenarios with dynamically changing requirements. In order to provide the expected performance and energy efficiency values in these situations, the use of heterogeneous hardware/software platforms at the edge has become widespread. However, these computing platforms still suffer from the lack of unified software-driven programming models to efficiently deploy multi-purpose hardware-accelerated solutions. In parallel, edge computing systems also face another huge challenge: operating under multiple conditions that were not taken into account during any of the design stages. Moreover, these conditions may change over time, forcing self-adaptation mechanisms to become a must. This paper presents an integrated architecture to exploit hardware-accelerated data-parallel models and transparent hardware/software multithreading. In particular, the proposed architecture leverages the ARTICo3framework and ReconOS to allow developers to select the most suitable programming model to deploy their edge computing applications onto run-time reconfigurable hardware devices. An evolvable hardware system is used as an additional architectural component during validation, providing support for continuous lifelong learning in smart edge computing scenarios. In particular, the proposed setup exhibits online learning capabilities that include learning by imitation from software-based reference algorithms. Experimental results show the benefits of the proposed approach, exposing different run-time tradeoffs (e.g., computing performance versus functional correctness of the evolved solutions), and highlighting the benefits of using scalable data-parallel models to perform circuit evolution under dynamically changing application scenarios. Alfonso Rodríguez 0002, Andrés Otero, Marco Platzner, Eduardo de la Torre |
IEEE Trans. Computers | 2 |
| 2021 | A Machine-Learning-Based Distributed System for Fault Diagnosis With Scalable Detection Quality in Industrial IoTabstractIn this article, a methodology based on machine learning for fault detection in continuous processes is presented. It aims to monitor fully distributed scenarios, such as the Tennessee Eastman process (TEP), selected as the use case of this work, where sensors are distributed throughout an industrial plant. A hybrid feature selection approach based on filters and wrappers, called hybrid Fisher wrapper method, is proposed to select the most representative sensors to get the highest detection quality for fault identification. The proposed methodology provides a complete design space of solutions differing in the sensing effort, the processing complexity, and the obtained detection quality. It constitutes an alternative to the typical scheme in Industry 4.0, where multiple distributed sensor systems collect and send data to a centralized cloud. Differently, the proposed technique follows a distributed approach, in which processing can be done eventually close to the sensors where data is generated, i.e., at the edge of the Internet of Things. This approach overcomes the bandwidth, privacy, and latency limitations that centralized approaches may suffer. The experimental results show that the proposed methodology provides TEP fault-detection solutions with state-of-the-art detection quality figures. In terms of latency, solutions obtained outperform in 37.5 times the implementation with the highest detection quality, using 1.99 times fewer features, on average. Also, the scalability of the framework provides a design space where the optimal implementation can be chosen according to the application needs. Rodrigo Marino, Cristian Wisultschew, Andrés Otero, José Manuel Lanza-Gutiérrez, Jorge Portilla, Eduardo de la Torre |
IEEE Internet Things J. | 3 |
| 2021 | Multi-grain reconfigurable and scalable overlays for hardware accelerator compositionabstractIn this paper, the reconfigurable nature of SRAM-based FPGAs is exploited to build a dynamically multi-grain reconfigurable and scalable overlay architecture. The composition of the overlay can be reconfigured on the fly to map applications with different requirements, change the size of the overlay to free up resources for other accelerators, or to deal with run-time variable computation demands. The overlay has been implemented by combining two different dynamic partial reconfiguration granularities. First, medium grain is used to compose the overlay by stitching together different processing elements, while fine grain reconfiguration is used to map applications onto the overlay by configuring the interconnections of the processing elements. The overlay has been coupled with a multiport memory and integrated into a System-on-Chip (SoC). Finally, a fully automated toolchain is proposed to transform code segments with one or two nested loops into the appropriate overlay configuration and the required control routines, enabling the transparent offloading of the computing-intensive parts of the application from the embedded SoC processor to hardware. The proposed overlay has been implemented in a Xilinx Zynq-7000 device and tested using the benchmark proposed in the CGRA-ME framework, obtaining up to 2× speed-up. Rafael Zamacola, Andrés Otero, Eduardo de la Torre |
J. Syst. Archit. | 2 |
| 2019 | Automated Tool and Runtime Support for Fine-Grain Reconfiguration in Highly Flexible Reconfigurable SystemsabstractDynamic partial reconfiguration significantly reduces reconfiguration times when offloading a partial design. However, there are occasions when fine-tuning a circuit would greatly benefit from quicker reconfiguration times. To that end, authors present an automated tool and runtime support to reconfigure LUT-based multiplexers and constants. In contrast to conventional multiplexers and constants, it is possible to modify these components without having a direct communication with the static system. Rafael Zamacola, Alberto García-Martínez, Javier Mora 0001, Andrés Otero, Eduardo de la Torre |
FCCM | 4 |
| 2015 | Sophisticated security verification on routing repaired balanced cell-based dual-rail logic against side channel analysisabstractConventional dual‐rail precharge logic suffers from difficult implementations of dual‐rail structure for obtaining strict compensation between the counterpart rails. As a light‐weight and high‐speed dual‐rail style, balanced cell‐based dual‐rail logic (BCDL) uses synchronised compound gates with global precharge signal to provide high resistance against differential power or electromagnetic analyses. BCDL can be realised from generic field programmable gate array (FPGA) design flows with constraints. However, routings still exist as concerns because of the deficient flexibility on routing control, which unfavourably results in bias between complementary nets in security‐sensitive parts. In this article, based on a routing repair technique, novel verifications towards routing effect are presented. An 8 bit simplified advanced encryption processing (AES)‐co‐processor is executed that is constructed on block random access memory (RAM)‐based BCDL in Xilinx Virtex‐5 FPGAs. Since imbalanced routing are major defects in BCDL, the authors can rule out other influences and fairly quantify the security variants. A series of asymptotic correlation electromagnetic (EM) analyses are launched towards a group of circuits with consecutive routing schemes to be able to verify routing impact on side channel analyses. After repairing the non‐identical routings, Mutual information analyses are executed to further validate the concrete security increase obtained from identical routing pairs in BCDL. Wei He 0015, Shivam Bhasin, Andrés Otero, Tarik Graba, Eduardo de la Torre, Jean-Luc Danger |
IET Inf. Secur. | 3 |
| 2014 | A dynamically adaptable bus architecture for trading-off among performance, consumption and dependability in Cyber-Physical SystemsabstractCyber-Physical Systems need to handle increasingly complex tasks, which additionally, may have variable operating conditions over time. Therefore, dynamic resource management to adapt the system to different needs is required. In this paper, a new bus-based architecture, called ARTICo3, which by means of Dynamic Partial Reconfiguration, allows the replication of hardware tasks to support module redundancy, multi-thread operation or dual-rail solutions for enhanced side-channel attack protection is presented. A configuration-aware data transaction unit permits data dispatching to more than one module in parallel, or provide coalesced data dispatching among different units to maximize the advantages of burst transactions. The selection of a given configuration is application independent but context-aware, which may be achieved by the combination of a multi-thread model similar to the CUDA kernel model specification, combined with a dynamic thread/task/kernel scheduler. A multi-kernel application for face recognition is used as an application example to show one scenario of the ARTICo3architecture. Juan Valverde, Alfonso Rodríguez 0002, Julio Camarero, Andrés Otero, Jorge Portilla, Eduardo de la Torre, Teresa Riesgo |
FPL | 4 |
| 2013 | A self-adaptive image processing application based on evolvable and scalable hardwareabstractEvolvable Hardware (EH) is a technique that consists of using reconfigurable hardware devices whose configuration is controlled by an Evolutionary Algorithm (EA). Our system consists of a fully-FPGA implemented scalable EH platform, where the Reconfigurable processing Core (RC) can adaptively increase or decrease in size. Figure 1 shows the architecture of the proposed System-on-Programmable-Chip (SoPC), consisting of a MicroBlaze processor responsible of controlling the whole system operation, a Reconfiguration Engine (RE), and a Reconfigurable processing Core which is able to change its size in both height and width. This system is used to implement image filters, which are generated autonomously thanks to the evolutionary process. The system is complemented with a camera that enables the usage of the platform for real time applications. Angel Gallego, Javier Mora 0001, Andrés Otero, Eduardo de la Torre, Teresa Riesgo |
FPL | 3 |
| 2013 | Self-Reconfigurable Evolvable Hardware System for Adaptive Image ProcessingabstractThis paper presents an evolvable hardware system, fully contained in an FPGA, which is capable of autonomously generating digital processing circuits, implemented on an array of processing elements (PEs). Candidate circuits are generated by an embedded evolutionary algorithm and implemented by means of dynamic partial reconfiguration, enabling evaluation in the final hardware. The PE array follows a systolic approach, and PEs do not contain extra logic such as path multiplexers or unused logic, so array performance is high. Hardware evaluation in the target device and the fast reconfiguration engine used yield smaller reconfiguration than evaluation times. This means that the complete evaluation cycle is faster than software-based approaches and previous evolvable digital systems. The selected application is digital image filtering and edge detection. The evolved filters yield better quality than classic linear and nonlinear filters using mean absolute error as standard comparison metric. Results do not only show better circuit adaptation to different noise types and intensities, but also a nondegrading filtering behavior. This means they may be run iteratively to enhance filtering quality. These properties are even kept for high noise levels (40 percent). The system as a whole is a step toward fully autonomous, adaptive systems. Rubén Salvador, Andrés Otero, Javier Mora 0001, Eduardo de la Torre, Teresa Riesgo, Lukás Sekanina |
IEEE Trans. Computers | 2 |
| 2012 | On the automatic integration of hardware accelerators into FPGA-based embedded systemsabstractThis paper proposes an automatic framework for the seamless integration of hardware accelerators, starting from an OpenMP-based application and an XML file describing the HW/SW partitioning. It extends a fully software architecture by generating and integrating the cores, along with the proper interfaces, and the code for scheduling and synchronization. Experimental results show that it is possible to validate different solutions only by varying the input code. Christian Pilato, Andrea Cazzaniga, Gianluca Durelli, Andrés Otero, Donatella Sciuto, Marco D. Santambrogio |
FPL | 4 |
| 2012 | Implementation techniques for evolvable HW systems: virtual VS. dynamic reconfigurationabstractAdaptive hardware requires some reconfiguration capabilities. FPGAs with native dynamic partial reconfiguration (DPR) support pose a dilemma for system designers: whether to use native DPR or to build a virtual reconfigurable circuit (VRC) on top of the FPGA which allows selecting alternative functions by a multiplexing scheme. This solution allows much faster reconfiguration, but with higher resource overhead. This paper discusses the advantages of both implementations for a 2D image processing matrix. Results show how higher operating frequency is obtained for the matrix using DPR. However, this is compensated in the VRC during evolution due to the comparatively negligible reconfiguration time. Regarding area, the DPR implementation consumes slightly more resources due to the reconfiguration engine, but adds further more capabilities to the system. Rubén Salvador, Andrés Otero, Javier Mora 0001, Eduardo de la Torre, Teresa Riesgo, Lukás Sekanina |
FPL | 2 |
| 2011 | Run-Time Scalable Architecture for Deblocking Filtering in H.264/AVC-SVC Video CodecsabstractSystems relying on fixed hardware components with a static level of parallelism can suffer from an under use of logical resources, since they have to be designed for the worst-case scenario. This problem is especially important in video applications due to the emergence of new flexible standards, like Scalable Video Coding (SVC), which offer several levels of scalability. In this paper, Dynamic and Partial Reconfiguration (DPR) of modern FPGAs is used to achieve run-time variable parallelism, by using scalable architectures where the size can be adapted at run-time. Based on this proposal, a scalable Deblocking Filter core (DF), compliant with the H.264/AVC and SVC standards has been designed. This scalable DF allows run-time addition or removal of computational units working in parallel. Scalability is offered together with a scalable parallelization strategy at the macro block (MB) level, such that when the size of the architecture changes, MB filtering order is modified accordingly. Andrés Otero, Eduardo de la Torre, Teresa Riesgo, Teresa Cervero, Sebastián López, Gustavo M. Callicó, Roberto Sarmiento |
FPL | 1 |
| 2011 | A novel scalable Deblocking Filter architecture for H.264/AVC and SVC video codecsabstractA highly parallel and scalable Deblocking Filter (DF) hardware architecture for H.264/AVC and SVC video codecs is presented in this paper. The proposed architecture mainly consists on a coarse grain systolic array obtained by replicating a unique and homogeneous Functional Unit (FU), in which a whole Deblocking-Filter unit is implemented. The proposal is also based on a novel macroblock-level parallelization strategy of the filtering algorithm which improves the final performance by exploiting specific data dependences. This way communication overhead is reduced and a more intensive parallelism in comparison with the existing state-of-the-art solutions is obtained. Furthermore, the architecture is completely flexible, since the level of parallelism can be changed, according to the application requirements. The design has been implemented in a Virtex-5 FPGA, and it allows filtering 4CIF (704 × 576 pixels @30 fps) video sequences in real-time at frequencies lower than 10.16 Mhz. Teresa Cervero, Andrés Otero, Sebastián López, Eduardo de la Torre, Gustavo M. Callicó, Roberto Sarmiento, Teresa Riesgo |
ICME | 2 |
| 2010 | A Modular Peripheral to Support Self-Reconfiguration in SoCsabstractIn this paper, a solution to support the run-time read back, relocation and replication of cores in embedded systems with dynamic and partial reconfiguration capabilities is presented. The proposal shows a peripheral structure that allows an easy integration and communication with the rest of the system, including an API to make the reconfiguration details to be more transparent to software applications. Differently to other proposals, all functionality is implemented in hardware, achieving a higher reconfiguration speed. In addition, different design decisions have been taken in order to increase the portability of the solution to existing and, possibly, future FPGAs. Finally, a use case is provided, which shows the features of this module applied to the run-time scaling of a hardware coprocessor. Andrés Otero, Angel Morales-Cas, Jorge Portilla, Eduardo de la Torre, Teresa Riesgo |
DSD | 1 |
| 2010 | Run-Time Scalable Systolic Coprocessors for Flexible Multimedia SoPCsabstractMultimedia Systems on Chip have high computational requirements, as well as significant flexibility demands. Flexibility can be related with the reusability of the cores in charge of the execution of computation-intensive tasks, but also with the run-time adaptation of these cores to the execution of time-variable tasks, or to changing system conditions. Among the run-time flexibility requirements of hardware IPs, functional scalability has been identified as an interesting feature. The proposal in this paper is to take advantage of the regularity and the high-processing capability of systolic arrays, to develop run-time functional scalable cores, making use of spatial scalability, by means of replicating and relocating basic processing elements of the array. The relocation process is performed using the dynamic-reconfiguration possibilities offered by commercial FPGAs. In this paper, an architectural template is proposed to develop systolic scalable coprocessors following this approach, together with its corresponding software drivers that may be executed within an embedded processor. In addition, a design flow is proposed to adapt the architectural template to different problems, together with some examples of scalable cores created following this design. This solution provides better results, regarding the reconfiguration time and the memory necessity overhead, compared with other dynamically scalable solutions. Andrés Otero, Eduardo de la Torre, Teresa Riesgo, Yana Esteves Krasteva |
FPL | 1 |