VLDB 2026 Research / reviewers in the wild / expert
José L. Núñez-Yáñez
dblp:99/3258 · also José Luis Núñez-Yáñez
· DBLP profile ↗
54ranked-venue papers
12as first author
7since 2021 · last 2026
0000-0002-5153-5481ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 12 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Software engineering, systems software and programming languages · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Reliability in Quantized Graph Neural Networks with Node-Wise Entropy-driven Temperature ScalingabstractGraph Neural Networks (GNNs) are one of the most powerful learning methods for graph-structured data and their quantization significantly reduces memory and computational requirements on edge devices. In this paper, we show that the quantization of node features, edge connections, and hidden representations degrades confidence calibration. To address this issue we propose a node-wise temperature scaling method that dynamically calibrates model confidence by aggregating entropy-based uncertainty from graph-structured data. Our approach combines self-entropy, neighborhood-entropy, and shortest-path distances to labeled nodes into a unified feature representation followed by a learnable transformation to compute temperature values for each node. We integrate and evaluate our approach using a dataflow hardware accelerator optimized for multi-precision GNN models, which supports efficient training and inference on-device. Our method significantly improves calibration by reducing Expected Calibration Error (ECE) and Negative Log-Likelihood (NLL) by up to 95% and 66%, respectively, across multiple datasets, without decreasing accuracy after quantization. The implementation is publicly available.1 Hadi Mousanejad Jeddi, José L. Núñez-Yáñez |
DATE | 2 |
| 2025 | SGRACE: Scalable Architecture for On-Device Inference and Training of Graph Attention and Convolutional NetworksabstractIn this article, we propose a hardware accelerator for on-device inference and training of deeply quantized graph convolutional networks (GCNs) and graph attention networks (GAT). The architecture unifies in a single dataflow both GAT and GCN support without breaking the matrix global formulation needed for high performance. In addition, it supports adaptive fixed-point precision (from 1- to 8-bit) and provides scalable performance through configurable hardware threads and compute units available in each thread. Each hardware thread employs a hierarchical dataflow combining fine- and coarse-grained components. The fine-grained dataflow streams words with a bit-width that depends on the selected precision. The coarse-grained dataflow connects aggregation and combination stages, incorporating an optional attention mechanism in GAT mode which is bypassed in GCN mode. In training mode, the hardware emulates multiple precisions (i.e., 1-/2-/4-/8-bit) to perform hardware-aware quantized training (HQT) for weights, features, attention coefficients, and the graph connectivity expressed in the normalized adjacency matrix. In inference mode, the architecture is optimized to support a single precision (i.e., 1- and 2-bit). The accelerator is mapped to a Zynq Ultrascale MPSOC and integrated within the Pytorch framework. Performance evaluation on Planetoid and Molecular graph datasets demonstrates over$100\times $acceleration for GCN and GAT compared to on-device CPU execution. The comparison with other FPGA, mobile GPU, and CPU hardware shows up to$50\times $better performance per joule. Results confirm HQT as an effective and efficient quantization strategy for graph neural network acceleration. José L. Núñez-Yáñez, Hadi Mousanejad Jeddi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Evaluation of Early-exit Strategies in Low-cost FPGA-based Binarized Neural NetworksabstractIn this paper, we investigate the application of early-exit strategies to quantized neural networks with binarized weights, mapped to low-cost FPGA SoC devices. The increasing complexity of network models means that hardware reuse and heterogeneous execution are needed and this opens the opportunity to evaluate the prediction confidence level early on. We apply the early-exit strategy to a network model suitable for ImageNet classification that combines weights with floating-point and binary arithmetic precision. The experiments show an improvement in inferred speed of around 20% using an early-exit network, compared with using a single primary neural network, with a negligible accuracy drop of 1.56%. Minxuan Kong, Kris Nikov, José L. Núñez-Yáñez |
DSD | 3 |
| 2022 | Analysis of Graph Processing in Reconfigurable Devices for Edge Computing ApplicationsabstractGraph processing is an area that has received significant attention in recent years due to the substantial expansion in industries relying on data analytics. Alongside the vital role of finding relations in social networks, graph processing is also widely used in transportation to find optimal routes and biological networks to analyse sequences. The main bottleneck in graph processing is irregular memory accesses rather than computation intensity. Since computational intensity is not a driving factor, we propose a method to perform graph processing at the edge more efficiently. We believe current cloud computing solutions are still very costly and have latency issues. The results demonstrate the benefits of a dedicated sparse graph processing algorithm compared with dense graph processing when analysing data with low density. As graph datasets grow exponentially, traversal algorithms such as breadth-first search (BFS), fundamental to many graph processing applications and metrics, become more costly to compute. Our work focuses on reviewing other implementations of breadth-first search algorithms designed for low power systems and proposing our solution that utilises advanced enhancements to achieve a significant performance boost up to 9.2x better performance in terms of MTEPS compared to other state-of-the-art solutions with a power usage of 2.32W. Kaan Olgu, Kris Nikov, José L. Núñez-Yáñez |
DSD | 3 |
| 2022 | A Low-complexity FPGA TDC based on a DSP Delay Line and a Wave Union LauncherabstractHigh-precision time-to-digital converters (TDCs) are key components for controlling quantum systems and FPGAs have gained popularity for this task thanks to their low-cost and flexibility compared with Application Specific Integrated Circuits (ASICs). This paper investigates a novel FPGA-based TDC architecture that combines a wave union launcher and delay lines constructed with DSP blocks. The configuration achieves a 8.07ps RMS resolution on a low-cost Zynq FPGA with a power usage of only 0.628W. The low power consumption is achieved thanks to a combination of operating frequency and logic resource usage that are lower than other methods, such as multi-chain DSP based TDCs and multi-chain CARRY4 based TDCs. Jiajun Lu, José L. Núñez-Yáñez |
DSD | 3 |
| 2022 | Lightweight asynchronous scheduling in heterogeneous reconfigurable systemsabstractThe trend for heterogeneous embedded systems is the integration of accelerators and general-purpose CPU cores on the same die. In these integrated architectures, like the Zynq UltraScale+ board (CPU+FPGA) that we target in this work, hardware support for shared memory and low-overhead synchronization between the accelerator and the CPU cores make the case for exploring strategies that exploit a tight collaboration between the CPUs and the accelerator. In this paper we propose a novel lightweight scheduling strategy, FastFit, targeted to FPGA accelerators, and a new scheduler based on it, named MultiFastFit, which asynchronously tackles heterogeneous systems comprised of a variety of CPU cores and FPGA IPs. Our strategy significantly reduces the overhead to automatically compute the near-optimal chunksizes when compared to a previous state-of-the-art auto-tuned approach, which makes our approach more suitable for fine-grained applications. Additionally, our scheduler MultiFastFit has been designed to enable the efficient co-execution of work among compute devices in such a way that all the devices are busy while minimizing the load unbalance. Our approaches have been evaluated using four benchmarks carefully tuned for the low-power UltraScale+ platform. Our experiments demonstrate that the FastFit strategy always finds the near-optimal FPGA chunksize for any device configuration at a reasonable cost, even for fine-grained and irregular applications, and that heterogeneous CPU+FPGA co-executions that exploit all the compute devices are usually faster and more energy efficient than the CPU-only and FPGA-only executions. We have also compared MultiFastFit with other state-of-the-art scheduling strategies, finding that it outperforms other auto-tuned approach up to 2x and it achieves similar results to manually-tuned schedulers without requiring an offline search of the ideal CPU-FPGA partition or FPGA chunk granularity. Andrés Rodríguez Moreno, Angeles G. Navarro, Kris Nikov, José L. Núñez-Yáñez, Ruben Gran Tejero, Darío Suárez Gracia, Rafael Asenjo |
J. Syst. Archit. | 4 |
| 2021 | Energy-efficient neural networks with near-threshold processors and hardware accelerators
José L. Núñez-Yáñez, Neil J. Howard |
J. Syst. Archit. | 1 |
| 2020 | A Streaming Dataflow Engine for Sparse Matrix-Vector Multiplication Using High-Level SynthesisabstractUsing high-level synthesis techniques, this paper proposes an adaptable high-performance streaming dataflow engine for sparse matrix dense vector multiplication (SpMV) suitable for embedded FPGAs. As the SpMV is a memory-bound algorithm, this engine combines the three concepts of loop pipelining, dataflow graph, and data streaming to utilize most of the memory bandwidth available to the FPGA. The main goal of this paper is to show that FPGAs can provide comparable performance for memory-bound applications to that of the corresponding CPUs and GPUs but with significantly less energy consumption. The experimental results indicate that the FPGA provides higher performance compared to that of embedded GPUs for small and medium-size matrices by an average factor of 3.25 whereas the embedded GPU is faster for larger size matrices by an average factor of 1.58. In addition, the FPGA implementation is more energy efficient for the range of considered matrices by an average factor of 8.9 compared to the embedded CPU and GPU. A case study based on adapting the proposed SpMV optimization to accelerate the support vector machine (SVM) algorithm, one of the successful classification techniques in the machine learning literature, justifies the benefits of utilizing the proposed FPGA-based SpMV compared to that of the embedded CPU and GPU. The experimental results show that the FPGA is faster by an average factor of 1.7 and consumes less energy by an average factor of 6.8 compared to the GPU. Mohammad Hosseinabady, José L. Núñez-Yáñez |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Parallel multiprocessing and scheduling on the heterogeneous Xeon+FPGA platform
Andrés Rodríguez Moreno, Angeles G. Navarro, Rafael Asenjo, Francisco Corbera, Ruben Gran Tejero, Darío Suárez Gracia, José L. Núñez-Yáñez |
J. Supercomput. | 7 |
| 2019 | Performance and Energy Efficiency Trade-Offs in Single-ISA Heterogeneous Multi-Processing for Parallel ApplicationsabstractThis work proposes a novel methodology to predict the optimal performance and energy efficiency trade-off configurations of parallel applications running on a two-cluster Heterogeneous Multi-Processing (HMP) system. we propose an analytic performance and power model that are generated offline using data measurements. These models are then used to estimate the whole configuration space to predict the application's performance and energy consumption. Then, we use these off-line predictions to choose Pareto-optimal configurations, which is the most efficient among all configurations for the given architecture and multi-threaded application. We validated our methodology on an ODROID XU3 board on several PARSEC and Phoronix Test Suite applications. Demetrios Coutinho, Kyriakos Georgiou, Kerstin Eder, José L. Núñez-Yáñez, Samuel Xavier de Souza |
VLSI-SoC | 4 |
| 2019 | Exploring heterogeneous scheduling for edge computing with CPU and FPGA MPSoCs
Andrés Rodríguez Moreno, Angeles G. Navarro, Rafael Asenjo, Francisco Corbera, Ruben Gran Tejero, Darío Suárez Gracia, José L. Núñez-Yáñez |
J. Syst. Archit. | 7 |
| 2019 | Energy Proportional Neural Network Inference with Adaptive Voltage and Frequency ScalingabstractThis research presents the extension and application of a voltage and frequency scaling framework called Elongate to a high-performance and reconfigurable binarized neural network. The neural network is created in the FPGA reconfigurable fabric and coupled to a multiprocessor host that controls the operational point to obtain energy proportionality. Elongate instruments a design netlist by inserting timing detectors to enable the exploitation of the operating margins of a device reliably. The elongated neural network is re-targeted to devices with different nominal operating voltages and fabricated with 28 nm (i.e., Zynq) and 16nm (i.e., Zynq Ultrascale) feature sizes showing the portability of the framework to advanced process nodes. New hardware and software components are created to support the 16nm fabric microarchitecture and a comparison in terms of power, energy and performance with the older 28 nm process is performed. The results show that Elongate can obtain new performance and energy points that are up to 86 percent better than nominal at the same level of classification accuracy. Trade-offs between energy and performance are also possible with a large dynamic range of valid working points available. The results also indicate that the built-in neural network robustness allows operation beyond the first point of error while maintaining the classification accuracy largely unaffected. José L. Núñez-Yáñez |
IEEE Trans. Computers | 1 |
| 2019 | Simultaneous multiprocessing in a software-defined heterogeneous FPGAabstractHeterogeneous chips that combine CPUs and FPGAs can distribute processing so that the algorithm tasks are mapped onto the most suitable processing element. New software-defined high-level design environments for these chips use general purpose languages such as C++ and OpenCL for hardware and interface generation without the need for register transfer language expertise. These advances in hardware compilers have resulted in significant increases in FPGA design productivity. In this paper, we investigate how to enhance an existing software-defined framework to reduce overheads and enable the utilization of all the available CPU cores in parallel with the FPGA hardware accelerators. Instead of selecting the best processing element for a task and simply offloading onto it, we introduce two schedulers, Dynamic and LogFit, which distribute the tasks among all the resources in an optimal manner. A new platform is created based on interrupts that removes spin-locks and allows the processing cores to sleep when not performing useful work. For a compute-intensive application, we obtained up to 45.56% more throughput and 17.89% less energy consumption when all devices of a Zynq-7000 SoC collaborate in the computation compared against FPGA-only execution. José L. Núñez-Yáñez, Sam Amiri, Mohammad Hosseinabady, Andrés Rodríguez Moreno, Rafael Asenjo, Angeles G. Navarro, Darío Suárez Gracia, Ruben Gran Tejero |
J. Supercomput. | 1 |
| 2019 | Correction to: Simultaneous multiprocessing in a software-defined heterogeneous FPGAabstractThe presentation of Table 2 was incorrect in the original article. The correct Table 2 is given below. The original article has been corr José L. Núñez-Yáñez, Sam Amiri, Mohammad Hosseinabady, Andrés Rodríguez Moreno, Rafael Asenjo, Angeles G. Navarro, Darío Suárez Gracia, Ruben Gran Tejero |
J. Supercomput. | 1 |
| 2018 | Multi-precision convolutional neural networks on heterogeneous hardwareabstractFully binarised convolutional neural networks (CNNs) deliver very high inference performance using single-bit weights and activations, together with XNOR type operators for the kernel convolutions. Current research shows that full binarisation results in a degradation of accuracy and different approaches to tackle this issue are being investigated such as using more complex models as accuracy reduces. This paper proposes an alternative based on a multi-precision CNN frame-work that combines a binarised and a floating point CNN in a pipeline configuration deployed on heterogeneous hardware. The binarised CNN is mapped onto an FPGA device and used to perform inference over the whole input set while the floating point network is mapped onto a CPU device and performs re-inference only when the classification confidence level is low. A light-weight confidence mechanism enables a flexible trade-off between accuracy and throughput. To demonstrate the concept, we choose a Zynq 7020 device as the hardware target and show that the multi-precision network is able to increase the BNN accuracy from 78.5% to 82.5% and the CPU inference speed from 29.68 to 90.82 images/sec. Sam Amiri, Mohammad Hosseinabady, Simon McIntosh-Smith, José L. Núñez-Yáñez |
DATE | 4 |
| 2018 | Workload Partitioning Strategy for Improved Parallelism on FPGA-CPU Heterogeneous ChipsabstractIn heterogeneous computing, efficient parallelism can be obtained if every device runs the same task on a different portion of the data set. This requires designing a scheduler which assigns data chunks to compute units proportional to their throughputs. For FPGA-CPU heterogeneous devices, to provide the best possible overall throughput, a scheduler should accurately evaluate the different performance behaviour of the compute devices. In this article, we propose a scheduler which initially detects the highest throughput each device can obtain for a specific application with negligible overhead and then partitions the dataset for improved performance. To demonstrate the efficiency of this method, we choose a Zynq UltraScale+ ZCU102 device as the hardware target and parallelise four applications showing that the developed scheduler can provide up to 94.06% of the throughput achievable at an ideal condition, with comparable power and energy consumption. Sam Amiri, Mohammad Hosseinabady, Andrés Rodríguez Moreno, Rafael Asenjo, Angeles G. Navarro, José L. Núñez-Yáñez |
FPL | 6 |
| 2018 | Dynamic Energy Management of FPGA Accelerators in Embedded SystemsabstractIn this article, we investigate how to utilise an Field-Programmable Gate Array (FPGA) in an embedded system to save energy. For this purpose, we study the energy efficiency of a hybrid FPGA-CPU device that can switch task execution between hardware and software with a focus on periodic tasks. To increase the applicability of this task switching, we also consider the voltage and frequency scaling (VFS) applied to the FPGA to reduce the system energy consumption. We show that in some cases, if the task’s period is higher than a specific level, the FPGA accelerator cannot reduce the energy consumption associated to the task and the software version is the most energy efficient option. We have applied the proposed techniques to a robot map creation algorithm as a case study which shows up to 38% energy reduction compared to the FPGA implementation. Overall, experimental results show up to 48% energy reduction by applying the proposed techniques at runtime on 13 individual tasks. Mohammad Hosseinabady, José L. Núñez-Yáñez |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | A systematic approach to design and optimise streaming applications on FPGA using high-level synthesisabstractThis paper proposes a systematic approach to help designers to optimise a given streaming application for FPGAs using High-Level Synthesis (HLS). The proposed technique specifically addresses the two main issues in a streaming application that are determining the exact amount of loop unrolling in the HLS code to increase the throughput and finding the optimum buffers' size to prevent deadlocks. To evaluate the proposed techniques two applications from the machine learning optimisation area are studied in the paper. These applications are Hessian-vector product and Conjugate Gradient (CG). The experimental results show up to 38× speed-up in throughput compared to the original streaming implementations provided by knowledgeable engineers using the dataflow, loop pipelining and FIFO channel related pragmas provided by the HLS tool. In addition, these applications show up to 2.98 GB/sec usage of memory bandwidth which is 93.1% of the total memory bandwidth available on the system. The source codes of the designs are available at https://github.com/Hosseinabady/csdfg-hls. Mohammad Hosseinabady, José L. Núñez-Yáñez |
FPL | 2 |
| 2016 | Energy efficient video fusion with heterogeneous CPU-FPGA devices
Alin Achim, Ian Hasler, Paul R. Hill, José L. Núñez-Yáñez |
DATE | 5 |
| 2016 | Energy Optimization in Commercial FPGAs with Voltage, Frequency and Logic ScalingabstractThis paper investigates the energy reductions possible in commercially available FPGAs configured to support voltage, frequency and logic scalability combined with power gating. Voltage and frequency scaling is based on in-situ detectors that allow the device to detect valid working voltage and frequency pairs at run-time while logic scalability is achieved with partial dynamic reconfiguration. The considered devices are FPGA-processor hybrids with independent power domains fabricated in 28 nm process nodes. The test case is based on a number of operational scenarios in which the FPGA side is loaded with a motion estimation core that can be configured with a variable number of execution units. The results demonstrate that voltage scalability reduces power by up to 60 percent compared with nominal voltage operation at the same frequency. The energy analysis show that the most energy efficiency core configuration depends on the performance requirements. A low performance scenario shows that serial computation is more energy efficient than the parallel configuration while the opposite is true when the performance requirements increase. An algorithm is proposed to combine effectively adaptive voltage/logic scaling and power gating in the proposed system and application. José L. Núñez-Yáñez, Mohammad Hosseinabady, Arash Beldachi |
IEEE Trans. Computers | 1 |
| 2015 | Evaluation of Hybrid Run-Time Power Models for the ARM Big.LITTLE ArchitectureabstractHeterogeneous processors, formed by binary compatible CPU cores with different microarchitectures, enable energy reductions by better matching processing capabilities and software application requirements. This new hardware platform requires novel techniques to manage power and energy to fully utilize its capabilities, particularly regarding the mapping of workloads to appropriate cores. In this paper we validate relevant published work related to power modelling for heterogeneous systems and propose a new approach for developing run-time power models that uses a hybrid set of physical predictors, performance events and CPU state information. We demonstrate the accuracy of this approach compared with the state-of-the-art and its applicability to energy aware scheduling. Our results are obtained on a commercially available platform built around the Samsung Exynos 5 Octa SoC, which features the ARM big.LITTLE heterogeneous architecture. Kris Nikov, José L. Núñez-Yáñez, Matthew Horsnell |
EUC | 2 |
| 2015 | Energy optimization of FPGA-based stream-oriented computing with power gatingabstractIn this paper, we propose a technique to improve the energy efficiency of FPGA devices by exploiting power gating during idle periods in streaming applications. The main idea is to shuffle idle periods during application execution so that the energy and timing overheads of turning the FPGA on and off can become acceptable. A key requirement is that fast FPGA-based accelerators are available and that the application follows a repetitive nature of execution. In this case, the accelerators work on a successive computing mode to accumulate the idle intervals in different iterations in order to make power gating feasible. Streaming on demand applications which are ubiquitous in embedded and portable devices are very good candidates to benefit from this technique. A case study is presented based on an MP3 player as the streaming application which shows up to 52.9% energy reduction. Mohammad Hosseinabady, José L. Núñez-Yáñez |
FPL | 2 |
| 2015 | Optimised OpenCL workgroup synthesis for hybrid ARM-FPGA devicesabstractThis paper presents a workgroup synthesis mechanism to compile an OpenCL kernel to FPGA-based accelerators embedded in a multi-core CPU system-on-a-chip (SoC). The OpenCL kernels considered in this paper exhibit regular data access patterns. Coping with the limited amount of internal memory in embedded FPGAs, the workgroup synthesis utilises a novel data access pattern formulation to describe the parallelism already provided by the OpenCL kernels. To provide an OpenCL framework prototype to validate the proposed technique, a source-to-source compiler that transforms the OpenCL kernel into C/C++ code is developed. Then vendor-specific high-level synthesis tools are used to convert the C/C++ code into the FPGA bitstream. Results based on popular real applications show up to 89.8% improvement in the execution time compared to other commercial FPGA OpenCL implementations. Mohammad Hosseinabady, José L. Núñez-Yáñez |
FPL | 2 |
| 2015 | Adaptive Voltage Scaling with In-Situ Detectors in Commercial FPGAsabstractThis paper investigates the limits of adaptive voltage scaling (AVS) applied to commercial FPGAs which do not specifically support voltage adaptation. An adaptive power architecture based on a modified design flow is created with in-situ detectors and dynamic reconfiguration of clock management resources. AVS is a power-saving technique that enables a device to regulate its own voltage and frequency based on workload, process and operating conditions in a closed-loop configuration. It results in significant improved energy profiles compared with dynamic voltage frequency scaling (DVFS) in which the device uses a number of pre-calculated valid working points. The results of deploying AVS in FPGAs with in-situ detectors shows power and energy savings exceeding 85 percent compared with nominal voltage operation at the same frequency. The in-situ detector approach compares favorably with critical path replication based on delay lines since it avoids the need of cumbersome and error-prone delay line calibration. José L. Núñez-Yáñez |
IEEE Trans. Computers | 1 |
| 2014 | Accurate power control and monitoring in ZYNQ boardsabstractZYNQ devices combine a dual-core ARM Cortex A9 processor and a FPGA fabric in the same die and in different power domains. In this paper we investigate the run-time power scaling capabilities of these devices using of-the-shelf boards and proposed accurate and fine-grained power control and monitoring techniques. The experimental results show that both software and hardware methods are possible and the right selection can yield different results in terms of control and monitoring speeds, accuracy of measurement, power consumption, and area overhead. The results also demonstrate that significant power margins are available in the FPGA device with different voltage configurations possible. This can be used to complement traditional voltage scaling techniques applied to the processor domain to obtain hybrid energy proportional computing platforms. Arash Beldachi, José L. Núñez-Yáñez |
FPL | 2 |
| 2014 | Run-time power gating in hybrid ARM-FPGA devicesabstractEnergy proportional computing (EPC) enables the allocation of energy to tasks depending on computational demands. Computing at full speed and then dynamically turning off modules when they are not required for a period of time can be used to obtain EPC and it is an alternative to voltage scaling techniques in which the computation is slowed down. This paper investigates the viability of physical power gating FPGA devices that incorporate a hardened processor in a different power domain. The run-time power gating approach is applied to Xilinx ZYNQ devices that incorporate a hardened Cortex A9 multi-processor. The paper demonstrates that power down followed by a full reconfiguration can be controlled by the embedded processor autonomously. The results show that the minimum time that the FPGA fabric must remain in power-off state for the technique to be energy efficient is in the order of milliseconds and up to 96% power reduction occurs when the fabric voltage is lowered below critical level. These results take into account the overheads of controlling the programmable voltage regulators interfaced to the FPGA and the overhead of the reconfiguration needed when the device must be returned to the active state. Mohammad Hosseinabady, José L. Núñez-Yáñez |
FPL | 2 |
| 2014 | Power modelling and capping for heterogeneous ARM/FPGA SoCsabstractLow-power processors and accelerators that were originally designed for the embedded systems market are emerging as building blocks for servers. Power capping has been actively explored as a technique to reduce the energy footprint of high-performance processors. The opportunities and limitations of power capping on the new low-power processor and accelerator ecosystem are less understood. This paper presents an efficient power capping and management infrastructure for heterogeneous SoCs based on hybrid ARM/FPGA designs. The infrastructure coordinates dynamic voltage and frequency scaling with task allocation on a customised Linux system for the Xilinx Zynq SoC. We present a compiler-assisted power model to guide voltage and frequency scaling, in conjunction with workload allocation between the ARM cores and the FPGA, under given power caps. The model achieves less than 5% estimation bias to mean power consumption. In an FFT case study, the proposed power capping schemes achieve on average 97.5% of the performance of the optimal execution and match the optimal execution in 87.5% of the cases, while always meeting power constraints. Yun Wu 0003, José L. Núñez-Yáñez, Roger F. Woods, Dimitrios S. Nikolopoulos |
FPT | 2 |
| 2014 | Joint video fusion and super resolution based on Markov random fieldsabstractIn this paper, a joint video fusion and super-resolution algorithm is proposed. The method addresses the problem of generating a high-resolution (HR) image from infrared (IR) and visible (VI) low-resolution (LR) images, in a Bayesian framework. In order to preserve better the discontinuities, a Generalized Gaussian Markov Random Field (MRF) is used to formulate the prior. Experimental results demonstrate that information from both visible and infrared bands is recovered from the LR frames in an effective way. José L. Núñez-Yáñez, Alin Achim |
ICIP | 2 |
| 2014 | Bayesian Video Super-Resolution With Heavy-Tailed Prior ModelsabstractIn this paper, we present a Bayesian-based superresolution algorithm that uses approximations of symmetric alpha-stable (SαS) Markov random fields as prior. The approximated SαS prior is used to perform maximum a posteriori (MAP) estimation for the high-resolution (HR) image reconstruction process. Compared with other state-of-the-art prior models, the proposed prior can better capture the heavy tails of the distribution of the HR image. Thus, the edges of the reconstructed HR image are preserved better in our method. As the corresponding energy function is nonconvex, the graduated nonconvexity method is used to solve the MAP estimation. Experiments confirm the better fit achieved by the proposed model to the actual data distribution and the consequent improvement in terms of visual quality over previously proposed super-resolution algorithms. José L. Núñez-Yáñez, Alin Achim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Video super-resolution using low rank matrix completionabstractIn this paper, a novel video super-resolution image reconstruction algorithm is proposed. We design a patch-based low rank matrix completion algorithm. The proposed algorithm addresses the problem of generating a high-resolution (HR) image from several low-resolution (LR) images, based on sparse representation and low-rank matrix completion. The approach represents observed LR frames in the form of sparse matrices and rearranges those frames into low dimensional constructions. Experimental results demonstrate that, high-frequency details in the super resolved images are recovered from the LR frames. The gains in terms of PSNR and SSIM are significant. José L. Núñez-Yáñez, Alin Achim |
ICIP | 2 |
| 2013 | Biophysically Accurate Foating Point Neuroprocessors for Reconfigurable LogicabstractThis paper presents a high-performance and biophysically accurate neuroprocessor architecture based on floating point arithmetic and compartmental modeling. It aims to overcome the limitations of traditional hardware neuron models that simplify the required arithmetic using fixed-point models. This can result in arbitrary loss of precision due to rounding errors and data truncation. On the other hand, a neuroprocessor based on a floating-point bio-inspired model, such as the one presented in this work, is able to capture additional cell properties and accurately mimic cellular behaviors required in many neuroscience experiments. The architecture is prototyped in reconfigurable logic obtaining a flexible and adaptable cell and network structure together with real time performance by using the available floating point hardware resources in parallel. The paper also demonstrates model scalability by combining the basic processor components that describe the soma, dendrite and synapse of organic cells to form more complex neuron structures. Yiwei Zhang 0002, Joe McGeehan, Edward Regan, José L. Núñez-Yáñez |
IEEE Trans. Computers | 5 |
| 2012 | Lossless video compression based on backward adaptive pixel-based fast motion estimation
Cedric Nishan Canagarajah, José L. Núñez-Yáñez, Raffaele Vitulli |
Signal Process. Image Commun. | 3 |
| 2012 | Video Super-Resolution Using Generalized Gaussian Markov Random FieldsabstractIn this letter, we present the first application of the Generalized Gaussian Markov Random Field (GGMRF) to the problem of video super-resolution. The GGMRF prior is employed to perform a maximum a posteriori (MAP) estimation of the desired high-resolution image. Compared with traditional prior models, the GGMRF can describe the distribution of the high-resolution image much better and can also preserve better the discontinuities (edges) of the original image. Previous work that used GGMRF for image restoration in which the temporal dependencies among video frames has not considered. Since the corresponding energy function is convex, gradient descent optimization techniques are used to solve the MAP estimation. Results show the super-resolved images using the GGMRF prior not only offers a good enhancement of visual quality, but also contain a significantly smaller amount of noise. José L. Núñez-Yáñez, Alin Achim |
IEEE Signal Process. Lett. | 2 |
| 2012 | Adaptive Voltage Scaling in a Dynamically Reconfigurable FPGA-Based PlatformabstractPower is an important issue limiting the applicability of Field Programmable Gate Arrays (FPGAs) since it is considered to be up to one order of magnitude higher than in ASICs. Recently, dynamic reconfiguration in FPGAs has emerged as a viable technique able to achieve power and cost reductions by time-multiplexing the required functionality at runtime. In this article, the applicability of Adaptive Voltage Scaling (AVS) to FPGAs is considered together with dynamic reconfiguration of logic and clock management resources to further improve the power profile of these devices. AVS is a popular power-saving technique in ASICs that enables a device to regulate its own voltage and frequency based on workload, fabrication, and operating conditions. The resulting processing platform exploits the available application-dependent timing margins to achieve a power reduction up to 85% operating at 0.58 volts compared with operating at a nominal voltage of 1 volt. The results also show that the energy requirements at 0.58 volts are aproximately five times lower compared with nominal voltage and this can be explained by the approximate cubic relation of static energy with voltage and the fact that the static component dominates power consumption in the considered FPGA devices. Atukem Nabina, José L. Núñez-Yáñez |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2012 | Cogeneration of Fast Motion Estimation Processors and Algorithms for Advanced Video CodingabstractThis paper presents a flexible and scalable motion estimation processor capable of supporting the processing requirements for high-definition (HD) video using the H.264 Advanced Video Codec, which is suited for FPGA implementation. Unlike most previous work, our core is optimized to execute all existing fast block matching algorithms, which we show to match or exceed the inter-frame prediction performance of traditional full-search approaches at the HD resolutions commonly in use today. Using our development tools, such algorithms can be described using a C-style syntax which is compiled into our custom instruction set. We show that different HD sequences exhibit different characteristics which necessitate a flexible and configurable solution when targeting embedded applications. This is supported in our core and toolset by allowing designers to modify the number of functional units to be instantiated. All processor instances remain binary compatible so recompilation of the motion estimation algorithm is not required. Due to this optimization process, it is possible to match the processing requirements of the selected motion estimation algorithm to the hardware microarchitecture leading to a very efficient implementation. José L. Núñez-Yáñez, Atukem Nabina, Eddie Hung, George Vafiadis |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | SystemC Architectural Transaction Level Modelling for Large NoCs
Mohammad Hosseinabady, José L. Núñez-Yáñez |
FDL | 2 |
| 2010 | Dynamic Reconfiguration Optimisation with Streaming Data DecompressionabstractThis paper presents a high performance reconfiguration controller enhanced with the use of streaming lossless decompression in its data path. Two reconfiguration controllers are designed, the first is a generic controller that utilises standard concepts such as Direct Memory Access, burst mode transfer of data and interrupts to maximise throughput. This controller is then improved by the inclusion of a streaming decompression engine optimised for the Internal Configuration Access Port (ICAP) interface. This new controller significantly improves the reconfiguration speed of the system and throughputs of up to 385 Mbytes/sec are recorded. As power and energy become very important constraints in the system design, an investigation of the overheads associated with the use of the reconfiguration controller are experimentally quantified and presented. Atukem Nabina, José L. Núñez-Yáñez |
FPL | 2 |
| 2010 | Effective modelling of large NoCs using SystemCabstractThe IEEE SystemC standard has been accepted as an effective high-level system modelling library among designers. However, in order to implement fast simulation models and to consider new ideas and requirements at system level, some enhancements and new features should be added to this standard. This is the reason why OSCI has proposed the TLM 1&2 libraries. In this work, we investigate a very fast but accurate and simple to use modelling methodology for systems which include a large number of modules such as complex Network-on-Chips (NoCs) which have many routers, network interfaces and processing cores. The proposed methodology implements all SystemC processes using normal functions and supports process activation simply by function calls. For this purpose, it utilises a SystemC method process that establishes concurrency and communication among the functions. The experimental results show an improvement of up to 98% in elaboration time and up to 90% in simulation time for small size NoCs. In addition, it can efficiently simulate large NoC with tens of thousands of nodes while traditional SystemC modelling based on threads struggles to simulate hundreds of nodes. Mohammad Hosseinabady, José L. Núñez-Yáñez |
ISCAS | 2 |
| 2009 | Run-time resource management in fault-tolerant network on reconfigurable chipsabstractThis paper investigates the challenges of run-time resource management in future coarse-grained network-on-reconfigurable-chips (NoRCs). Run-time reconfiguration is a key feature expected in future processing systems which must support multiple applications whose processing requirements are not known at design time. This paper investigates a stochastic routing algorithm in a NoC-based system with dynamically reconfigurable tiles, able to cope with the dynamic behaviour of run-time task mapping. Experimental results show the efficiency of the proposed stochastic task mapping. Mohammad Hosseinabady, José L. Núñez-Yáñez |
FPL | 2 |
| 2009 | A toolset for the analysis and optimization of motion estimation algorithms and processorsabstractThis paper presents a reconfigurable processor designed to execute user-defined block-matching motion estimation algorithms, and a toolset for the design of such algorithms and for the configuration of the processor. The toolset enables the exploration of the processor's design space in order to find an optimal configuration depending on the target application. The use of the toolset to test different configurations for different kinds of video sequences is illustrated. Experimental results show the benefits and cost of certain optimizations in the motion estimation process, and that fast block-matching search algorithms can outperform full search algorithms commonly used in hardware implementations. The usefulness of the toolset in exploring the configuration space is also shown. Trevor Spiteri, George Vafiadis, José L. Núñez-Yáñez |
FPL | 3 |
| 2009 | A biophysically accurate floating point somatic neuroprocessorabstractBiophysically accurate neuron models have emerged as a very useful tool for neuroscience research. These models are based on solving differential equations that govern membrane potentials and spike generation. The level of detail that needs to be presented in the model to accurately emulate the behaviour of an organic cell is still an open question, although the timing of the spikes is considered to convey essential information. Models targeting hardware are traditionally based on fixed point implementations and low precision algorithms which incur a significant loss of information. This, in turn, could affect the functionality of a bioelectronic neuroprocessor in an undefined way. In this paper, a 32-bit floating point reconfigurable somatic neuroprocessor is presented targeting an FPGA device for real-time processing. For each individual neuron, the dynamics of ionic channels are described by a set of first order kinetic equations. A dedicated CORDIC unit is developed to solve the nonlinear functions that regulate spike generation. The results have been verified using an experimental setup that combines an FPGA device and a digital-to-analogue converter. Yiwei Zhang 0002, José L. Núñez-Yáñez, Joe McGeehan, Edward Regan |
FPL | 2 |
| 2009 | Adaptive stochastic routing in fault-tolerant on-chip networksabstractDue to shrinking transistor geometries, on-chip circuits are becoming vulnerable to errors, but at the same time on-chip networks are required to provide reliable services over unreliable physical interconnects. A connection oriented stochastic routing (COSR) algorithm has been used on one NoC platform that provides excellent fault-tolerance and dynamic reconfiguration capability. A probability model has been built to analyze the COSR algorithm. According to the model, the performance may be improved by implementing a self learning mechanism in each router. Thus a new adaptive stochastic routing (ASR) algorithm is proposed whereby each router learns the network status from acknowledgement flits and stores the outcomes in a routing table. Simulation of both algorithms reveals that the ASR algorithm shows a higher path reservation success rate and a larger maximal accepted traffic than the COSR algorithm. The simulations also show that the learning procedures are accurate and that both algorithms are fault-tolerant to intermittent/permanent errors. Wei Song 0002, Doug A. Edwards, José L. Núñez-Yáñez, Sohini Dasgupta |
NOCS | 3 |
| 2009 | Backward Adaptive Pixel-based Fast Predictive Motion EstimationabstractThis letter presents a novel backward adaptive pixel-based fast predictive motion estimation (BAPME) scheme for lossless video compression. Unlike the widely used block-matching motion estimation techniques, this method predicts the motion on a pixel-by-pixel basis by comparing a group of past observed pixels in two adjacent frames, eliminating the need of transmitting side information. Combined with prediction and a fast search technique, the proposed algorithm achieves better entropy results and significant reduction in computation than pixel-based full search for a set of standard test sequences. Experimental results also suggest that BAPME is superior to block-based full search in terms of speed and zero-order entropy. Cedric Nishan Canagarajah, José L. Núñez-Yáñez |
IEEE Signal Process. Lett. | 3 |
| 2008 | Fault-tolerant dynamically reconfigurable NoC-based SoCabstractThis paper proposes a network-on-chip (NoC)-based dynamically reconfigurable platform which can perform multiple applications, simultaneously. A tile attached to a router in the NoC consists of a core container which can host a core permanently or temporarily. The tile also has a hardwired controller and a cache like memory to control the hosted cores. A core, which runs a task, may be described by a bitstream (called hardware core) or a programme code (called software core). Because of the dynamic behaviour of the proposed platform, using task identifier, a stochastic dynamic routing algorithm will find (or map) the task in the platform. Because of using the task identifier in routing algorithm and the reconfigurability of tiles, the proposed platform can tolerate probable faults. The proposed SoC architecture is easily able to run new protocols and tasks. Our results show that, the proposed platform follows the user interests such that runs tasks with higher temporal locality much faster than the tasks with lower temporal locality. Mohammad Hosseinabady, José L. Núñez-Yáñez |
ASAP | 2 |
| 2008 | Power/Area Analysis of a FPGA-Based Open-Source Processor using Partial Dynamic ReconfigurationabstractThis paper explores the utilization of run-time partial dynamic reconfiguration in the LEON3 open-source soft core processor, which is a highly configurable SPARC (scalable processor architecture) V8 instruction set processor. The work explores the possibilities of sharing different arithmetic functions tightly coupled to the integer pipeline and mapped to the same silicon area, saving power consumption and area utilisation. The same strategy can be used to extend the instruction set architecture of the processor with new instructions that are optimized for DSP applications. The logic necessary to support these instructions could then be swapped as demanded by the application. Izhar Zaidi, Atukem Nabina, Cedric Nishan Canagarajah, José L. Núñez-Yáñez |
DSD | 4 |
| 2008 | A configurable and programmable motion estimation processor for the H.264 video codecabstractThis work presents a programmable, configurable motion estimation processor for the H.264 video coding standard, capable of handling the processing requirements of high definition (HD) video and suitable for FPGA implementation. The programmable aspect of the processor follows the ASIP (Application Specific Instruction set Processor) approach with a instruction set targeted to accelerating block matching motion estimation algorithms. Configurability relates to the ability to optimize the microarchitecture for the selected algorithm and performance requirements through varying the number and type of execution units at compile time. José L. Núñez-Yáñez, Eddie Hung, Vassilios A. Chouliaras |
FPL | 1 |
| 2008 | Evaluating dynamic partial reconfiguration in the integer pipeline of a FPGA-based opensource processorabstractThis work explores the potential of sharing different arithmetic hardware operators tightly coupled to the integer pipeline of the open-source LEON3 processor. The idea is to map these modules to the same silicon area saving power consumption and area utilisation. The same strategy can be used to extend the architecture of processors optimized for applications with specific energy constraints. The proposed platform serves as a guideline to illustrate gains obtained through partial reconfiguration that need to adapt to changing standards and protocols with a limited number of resources. Izhar Zaidi, Atukem Nabina, Cedric Nishan Canagarajah, José L. Núñez-Yáñez |
FPL | 4 |
| 2008 | Customization of an embedded RISC CPU with SIMD extensions for video encoding: A case study
Vassilios A. Chouliaras, Vincent M. Dwyer, Shahrukh Agha, José L. Núñez-Yáñez, Dionysios I. Reisis, Konstantinos Nakos, Konstantinos Manolopoulos |
Integr. | 4 |
| 2008 | A Novel Delta Sigma Control System Processor and Its VLSI ImplementationabstractThis paper describes a novel control system processor architecture based on DeltaSigma modulation known as the DeltaSigma -CSP. The DeltaSigma -CSP utilizes 1-bit processing which is a new concept in digital control applications with the direct benefit of making multi-bit multiplication operations redundant. A simple conditional-negate-and-add (CNA) unit is instead used for operations in control law implementations. For this reason, the proposed processor has a very small silicon footprint and runs at very high frequencies making it ideal for high-sampling rate, real-time control applications. A number of DeltaSigma -CSP configurations have been implemented as VLSI hard macros in a high-performance 0.13-mum CMOS process and a particular configuration achieved a post-route operating frequency of 355 MHz resulting in a 2.17 MHz sampling rate for a fourth-order control law implementation. Additional results prove that the DeltaSigma -CSP compares very favorably, in terms of silicon area and sampling rates, to two other specialized digital control processing systems, including direct, hardwired implementation of control laws; at the same time, it substantially outperforms software implementations of control laws running on very wide, general-purpose VLIW architectures. Xiaofeng Wu 0001, Vassilios A. Chouliaras, José L. Núñez-Yáñez, Roger Goodall |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Dynamic Voltage Scaling in a FPGA-based System-on-ChipabstractThis paper presents a DVS (Dynamic Voltage Scaling) enabled SoC (System-on-Chip) processing platform based on the Leon3 open-source processor and dynamically reconfigurable clock synthesis technology available in Virtex-4 Xilinx FPGAs. A special DVS monitor unit maintains correct operation of the processor core at a given voltage by tracking the behavior of an internal delay line and stopping the processor clock through a digital clock management (DCM) macroblock when a timing error is about to occur. Upon detection of a new valid working point the DVS monitor unit reconfigures the main DCM to synthesize a new frequency-adjusted CPU clock signal and reactivates the processor. The energy savings and operation range of the technology are evaluated in the context of video coding applications by executing different motion estimation kernels. José L. Núñez-Yáñez, Vassilios A. Chouliaras, Jiri Gaisler |
FPL | 1 |
| 2007 | On the Hardware Reduction of z-Datapath of Vectoring CORDICabstractIn this article we present a novel design of a hardware optimal vectoring CORDIC processor. We present a mathematical theory to show that using bipolar binary notation it is possible to eliminate all the arithmetic computations required along the z-datapath. Using this technique it is possible to achieve three and 1.5 times reduction in the number of registers and adder respectively compared to classical CORDIC. Following this, a 16-bit vectoring CORDIC is designed for the application in Synchronizer for IEEE 802.11a standard. The total area and dynamic power consumption of the processor is 0.14 mm2and 700μW respectively when synthesized in 0.18μm CMOS library which shows its effectiveness as a low-area low-power processor. R. Stapenhurst, Koushik Maharatna, Jimson Mathew, José L. Núñez-Yáñez, Dhiraj K. Pradhan |
ISCAS | 4 |
| 2005 | A Thread and Data-Parallel MPEG-4 Video Encoder for a System-On-Chip MultiprocessorabstractWe studied the dynamic instruction count reduction for a single-thread, vectorized and a multithreaded, nonvectorized, MPEG-4 video encoder. Results indicate a maximum improvement of the order of 88% for 22 CPU contexts for the multithreaded case whereas the single-thread, vectorized version demonstrates an 85% improvement for a vector register file length of 24 bytes, over the scalar case. We present VLSI macrocells of a vector accelerator implementing a subset of the MPEG-4 vector ISA and a 2-way, parametric, bus-based, cache coherent, SoC multiprocessor. Tom R. Jacobs, José L. Núñez-Yáñez |
ASAP | 2 |
| 2005 | Design and Implementation of a High-Performance and Silicon Efficient Arithmetic Coding Accelerator for the H.264 Advanced Video CodecabstractA high performance and silicon efficient hardware architecture for binary arithmetic coding (BAC) acceleration is presented and its application to entropy coding in the context of the H.264 video compressor standard described. The proposed hardware architecture remains bit compatible with the software implementation used in the H.264 ITU standard. The renormalization sequence that maintains the state variables in the appropriate range has been rewritten in order to enable a data independent throughput in hardware of 1 symbol per clock cycle. The instruction set extensions required to be implemented as part of the ISA of a controlling RISC are proposed. Finally, ASIC and FPGA implementations are obtained and the performance and complexity compared with recent implementations of the well-known MQ-coder reported. José L. Núñez-Yáñez, Vassilios A. Chouliaras |
ASAP | 1 |
| 2005 | A Configurable Statistical Lossless Compression Core Based on Variable Order Markov Modeling and Arithmetic CodingabstractThis paper presents a practical realization in hardware of the concepts of variable order Markov modeling using multisymbol alphabets and arithmetic coding for lossless compression of universal data. This type of statistical coding algorithm has long been regarded as being able to deliver very high compression ratios close to the information content of the source data. However, their high computational complexity has limited their practical application in embedded environments such as in mobile computing and wireless communications. In this paper, a hardware amenable algorithm named PPMH and based on these principles has been developed and its architecture and implementation detailed. This novel lossless compression core offers innovative solutions to the computational issues in both stages of modeling and coding and delivers high compression efficiency and throughput. The configurability features of the core allow efficient use of the embedded SRAM present in modern FPGA technologies where memory resources range from a few kilobits to several megabits per device family. The core has been targeted to the Altera Stratix FPGA family and performance, coding efficiency, and complexity measured for different memory configurations. José L. Núñez-Yáñez, Vassilios A. Chouliaras |
IEEE Trans. Computers | 1 |