EDBT 2026 Demo / reviewers in the wild / expert
Alberto A. Del Barrio
dblp:99/695 · also Alberto A. Del Barrio García, Alberto Antonio Del Barrio García
· DBLP profile ↗
30ranked-venue papers
13as first author
9since 2021 · last 2024
0000-0002-6769-1200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 13 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Theory of computation · 3 · 2 since 2021Artificial intelligence and machine learning · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Square Root Unit with Minimum Iterations for Posit ArithmeticabstractIn this paper, we introduce a novel implementation of a square root algorithm specifically tailored for posit arithmetic. Unlike traditional methods, the proposed approach capitalizes on the inherent flexibility of posits, which lack fixed-length fields, to optimize square root computations. By accurately estimating the minimum number of required fraction bits, our algorithm substantially reduces the recurrence iterations without sacrificing accuracy. Implemented across standard 16-bit, 32-bit, and 64-bit posit formats, our units showcase a significant latency reduction in different applications with only a marginal increase in resource utilization. Comparative analysis against previous pipelined designs underscores the area efficiency of our proposed solutions. This research significantly contributes to the advancement of posit-based arithmetic units, presenting promising opportunities for improving computational system efficiency. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ARITH | 2 |
| 2024 | 1-D Spatial Attention in Binarized Convolutional Neural NetworksabstractThis paper proposes a structure called SPBNet for enhancing binarized convolutional neural networks (BCNNs) using a low-cost 1-D spatial attention structure. Attention blocks can compensate for the performance drop in BCNNs. However, the hardware overhead of complex attention blocks can be a significant burden in BCNNs. The proposed attention block consists of low-cost 1-D height-wise and width-wise 1-D convolutions, It has the attention bias to adjust the effects of attended features in ×0.5−×1.5. In experiments, the proposed block used in ResNet18-based BCNNs improves Top-1 accuracy up to 2.7% over a baseline ReActNet on the CIFAR100 dataset. Notably, without using teacher-student training, the proposed structure can show comparable performance as the baseline ReActNetA using teacher-student training. Hyun Jin Kim 0001, Jungwoo Shin, Alberto A. Del Barrio |
ICASSP | 4 |
| 2024 | Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific ComputingabstractThe accuracy requirements in many scientific computing workloads result in the use of double-precision floating-point arithmetic in the execution kernels. Nevertheless, emerging real-number representations, such as posit arithmetic, show promise in delivering even higher accuracy in such computations. In this work, we explore the native use of 64-bit posits in a series of numerical benchmarks and compare their timing performance, accuracy and hardware cost to IEEE 754 doubles. In addition, we also study the conjugate gradient method for numerically solving systems of linear equations in real-world applications. For this, we extend the PERCIVAL RISC-V core and the Xposit custom RISC-V extension with posit64 and quire operations. Results show that posit64 can obtain up to 4 orders of magnitude lower mean square error than doubles. This leads to a reduction in the number of iterations required for convergence in some iterative solvers. However, leveraging the quire accumulator register can limit the order of some operations such as matrix multiplications. Furthermore, detailed FPGA and ASIC synthesis results highlight the significant hardware cost of 64-bit posit arithmetic and quire. Despite this, the large accuracy improvements achieved with the same memory bandwidth suggest that posit arithmetic may provide a potential alternative representation for scientific computing. David Mallasén, Alberto A. Del Barrio, Manuel Prieto 0001 |
IEEE Trans. Computers | 2 |
| 2023 | A Suite of Division Algorithms for Posit ArithmeticabstractPosit ™ arithmetic is a promising alternative to IEEE 754 floating-point arithmetic due to its higher accuracy, larger dynamic range, and bitwise compatibility. While posit arithmetic has been well studied for basic arithmetic operations, division has received little attention. This paper proposes multiple divider designs for posit arithmetic based on digit recurrence and iterative approximation, and evaluates their performance. ASIC synthesis results show that the proposed designs significantly reduce hardware requirements for 32-bit division units compared to previous works by 1.14 × in area, 1.11× in power, and 1.04×in datapath delay. Moreover, the paper introduces an approximate logarithmic posit division that achieves an 8.8×reduction in area and 29×reduction in energy consumption with negligible degradation of the final results, making it suitable for error-tolerant applications. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ASAP | 2 |
| 2023 | PERCIVAL: Deploying Posits and Quire Arithmetic into the CVA6 RISC-V CoreabstractRepresenting and operating on real numbers in a microprocessor presents unique challenges not encountered with the set of integers. Working with real numbers introduces additional concepts such as precision, that is, the error made between the number with which we want to operate and the approximation that we can represent in a finite number of bits. Currently, the universally extended way of representing the set of real numbers is using floating-point numbers defined by the IEEE 754 standard. This format presents a series of difficulties, such as the different rounding schemes, reproducibility problems depending on the implementation, a multitude of ways to represent Not a Numbers (NaNs) or the existence of plus and minus zero. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
CF | 3 |
| 2023 | Generating Posit-Based Accelerators With High-Level SynthesisabstractRecently, the posit number system has demonstrated a higher accuracy over standard floating-point arithmetic for many scientific applications. However, when it comes to implementing accelerators for these applications, the tool support for this arithmetic format is still missing, especially during the step. In this paper, we incorporate the posit data type into the high-level synthesis (HLS) design process, so that we can generate the implementation directly from a given behavioral specification, but using posit numbers instead of the classical floating-point notations. Our evaluations show that, even if posit-based circuits require more area than their floating-point counterparts, they offer higher accuracy when using the same bitwidth. For example, using posit arithmetic can reduce computation errors by about two orders of magnitude when compared to using standard floating-point numbers. Our approach also includes an alternative to mitigate the high overheads of the posits and broadening the potential use of this format. We also propose a hybrid scheme that uses posit numbers only in the private local memory, while the accelerator operates in the classic floating-point notation. This solution is useful when the designers want to optimize local memories and data transfers, but still use legacy high-level synthesis (HLS) tools that only support traditional floating-point notations. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Christian Pilato |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | PERCIVAL: Open-Source Posit RISC-V Core With Quire CapabilityabstractPresents the front cover, title page, cover page, or splash screen of the proceedings record. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
ARITH | 3 |
| 2021 | Energy-Efficient MAC Units for Fused Posit ArithmeticabstractPosit arithmetic is an alternative format to the standard IEEE 754 for floating-point numbers that claims to provide compelling advantages over floats, including higher accuracy, larger dynamic range, or bitwise compatibility across systems. The interest in the design of arithmetic units for this novel format has increased in the last few years. However, while multiple designs for posit adder and multiplier have been developed recently in the literature, fused units for posit arithmetic are still in the early stages of research. Moreover, due to the large size of accumulators needed in fused operations, the few fused posit units proposed so far still require many hardware resources. In order to contribute to the development of the posit number format, and facilitate its use in applications such as deep learning, this paper presents several designs of energy-efficient posit multiply- accumulate (MAC) units with support for standard quire format. Concretely, the proposed designs are capable of computing fused dot products of large vectors without accuracy drop, while consuming less energy than previous implementations. Experiments show that, compared to previous implementations, the proposed designs consume up to 75.49%, 88.45% and 83.43% less energy and are 73.18%, 87.36% and 83.00% faster for 8, 16 and 32 bitwidths, with an additional area of only 4.97%, 7.44% and 4.24%, respectively. Raul Murillo 0001, David Mallasén, Alberto A. Del Barrio, Guillermo Botella Juan |
ICCD | 3 |
| 2021 | First experiences of teaching quantum computing
Ginés Carrascal, Alberto A. Del Barrio, Guillermo Botella Juan |
J. Supercomput. | 2 |
| 2020 | Customized Posit Adders and Multipliers using the FloPoCo Core GeneratorabstractThe posit number system, which is proposed as a replacement of IEEE floating-point numbers, is in the spotlight of Arithmetic research due to the recent breakthroughs. This format claims to provide more accurate results with the same bitwidth than standard floating point, but the run-time variability during the detection of the posit fields involves a hardware design challenge. In this work, we propose parameterized designs for multiple posit functional units, including addition and multiplication, and integrate them as templates of the FloPoCo framework. The integration of the proposed algorithms within FloPoCo can provide synthesizable VHDL code for posit arithmetic of any possible configuration 〈n, es〉. Experiments show an improvement in terms of area and energy with respect to state-of-the-art works up to 35.9% and 30.8%, respectively. Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan |
ISCAS | 2 |
| 2020 | HEVC optimization based on human perception for real-time environments
David Guillermo Fernández, Guillermo Botella Juan, Alberto A. Del Barrio, Carlos García 0001, Manuel Prieto 0001, Christos Grecos |
Multim. Tools Appl. | 3 |
| 2019 | A Cost-Efficient Iterative Truncated Logarithmic Multiplication for Convolutional Neural NetworksabstractThis paper proposes a cost-efficient approximate logarithmic multiplication for convolutional neural networks (CNNs), where two truncated logarithmic multipliers are connected for error correction. The proposed iterative logarithmic multiplication achieves low and unbiased average error while the hardware cost is significantly reduced by utilizing the truncated Mitchell multiplier and approximating error terms from the first stage. The proposed design has error characteristics that are suitable for neural network inferences, and the experiments on contemporary CNNs show that the proposed multiplier does not cause significant degradation on accuracy compared to exact multiplication. Hyun Jin Kim 0001, Min Soo Kim 0003, Alberto A. Del Barrio, Nader Bagherzadeh |
ARITH | 3 |
| 2019 | Design of Power-Efficient FPGA Convolutional Cores with Approximate Log Multiplier
Leonardo Tavares Oliveira, Min Soo Kim 0003, Alberto A. Del Barrio, Nader Bagherzadeh, Ricardo Menotti |
ESANN | 3 |
| 2019 | Efficient Mitchell's Approximate Log Multipliers for Convolutional Neural NetworksabstractThis paper proposes energy-efficient approximate multipliers based on the Mitchell’s log multiplication, optimized for performing inferences on convolutional neural networks (CNN). Various design techniques are applied to the log multiplier, including a fully-parallel LOD, efficient shift amount calculation, and exact zero computation. Additionally, the truncation of the operands is studied to create the customizable log multiplier that further reduces energy consumption. The paper also proposes using the one’s complements to handle negative numbers, as an approximation of the two’s complements that had been used in the prior works. The viability of the proposed designs is supported by the detailed formal analysis as well as the experimental results on CNNs. The experiments also provide insights into the effect of approximate multiplication in CNNs, identifying the importance of minimizing the range of error.The proposed customizable design at $w$w = 8 saves up to 88 percent energy compared to the exact fixed-point multiplier at 32 bits with just a performance degradation of 0.2 percent for the ImageNet ILSVRC2012 dataset. Min Soo Kim 0003, Alberto A. Del Barrio, Leonardo Tavares Oliveira, Román Hermida, Nader Bagherzadeh |
IEEE Trans. Computers | 2 |
| 2018 | Low-power implementation of Mitchell's approximate logarithmic multiplication for convolutional neural networksabstractThis paper proposes a low-power implementation of the approximate logarithmic multiplier to improve the power consumption of convolutional neural networks for image classification, taking advantage of its intrinsic tolerance to error. The approximate logarithmic multiplier converts multiplications to additions by taking approximate logarithm and achieves significant improvement in power and area while having low worst-case error, which makes it suitable for neural network computation. Our proposed design shows a significant improvement in terms of power and area over the previous work that applied logarithmic multiplication to neural networks, reducing power up to 76.6% compared to exact fixed-point multiplication, while maintaining comparable prediction accuracy in convolutional neural networks for MNIST and CIFAR10 datasets. Min Soo Kim 0003, Alberto A. Del Barrio, Román Hermida, Nader Bagherzadeh |
ASP-DAC | 2 |
| 2018 | Intra-Steganography: Hiding Data in High-Resolution VideosabstractSteganography is the art of hiding information within a file like an image or a video. The embedded message can then be used either to transmit some secret information or to protect the content of the file. Due to the increase of the resolutions to provide higher quality in videos, it is critical to comply with the latest video standard, namely: the High Efficiency Video Coding (HEVC), which allows reducing the size of the file to be transmitted. Thus, in this paper we propose an HEVC-compliant method to hide and retrieve information in high-resolution videos. The procedure is based on modifying the luminance of certain blocks. Nevertheless, this must be carefully done, as the HEVC standard is a powerful attack in itself, since it compresses 50% the size of the video on average and the embedded information may disappear. In this paper there is a study evaluating a proper spot to embed the information. Results show that it is possible to retrieve all the information while maintaining the quality of the video after embedding the message. Several tests have been run with the reference software HM-16.2 as well as the real-time encoder x265, and the SSIM and PSNR values are coherent with these theses. David Rodríguez 0002, Alberto A. Del Barrio, Guillermo Botella Juan, David Cuesta |
DS-RT | 2 |
| 2018 | Fast and effective CU size decision based on spatial and temporal homogeneity detection
David Guillermo Fernández, Alberto A. Del Barrio, Guillermo Botella Juan, Carlos García 0001 |
Multim. Tools Appl. | 2 |
| 2017 | A slack-based approach to efficiently deploy radix 8 booth multipliersabstractIn 1951 A. Booth published his algorithm to efficiently multiply signed numbers. Since the appearance of such algorithm, it has been widely accepted that radix 4-based Booth multipliers are the most efficient. They allow the height of the multiplier to be halved, at the expense of a simple recoding that consists of just shifts and negations. Theoretically, higher radix should produce even larger reductions, especially in terms of area and power, but the recoding process is much more complex. Notably, in the case of radix 8 it is necessary to compute 3X, X being the multiplicand. In order to avoid the penalty due to this calculation, we propose decoupling it from the product and considering 3X as an extra operation within the application's Dataflow Graph (DFG). Experiments show that typically there is enough slack in the DFGs to do this without degrading the performance of the circuit, which permits the efficient deployment of radix 8 multipliers that do not calculate the 3X multiple. Results show that our approach is 10% and 17% faster than radix 4 and radix 8 Booth based implementations, respectively, and 12% and 10% more energy efficient in terms of Energy Delay Product. Alberto A. Del Barrio, Román Hermida |
DATE | 1 |
| 2016 | A Partial Carry-Save On-the-Fly Correction Multispeculative MultiplierabstractFunctional Units that are designed to receive inputs and produce outputs using a non-redundant format typically exhibit an inferior performance. In order to overcome this limitation, the carry-save and partial carry-save formats have been proposed. Both approaches are very suitable when implementing addition trees. Nevertheless, if there are multiplications in the datapath, the inputs to the multiplier must be reduced to a non-redundant form, to avoid applying the distributive property. In this paper we present a multiplier able to receive two numbers in partial carry-save format, and produce a result in partial carry-save format as well. This is done by modifying the Booth encoder and leveraging the generate and propagate group signals that are available because of the partial carry-save format. Hence, this can allow to fully implement datapaths without additional penalty cycles due to reductions to non-redundant forms. Experiments show that our proposed multiplier has 15 percent shorter delay with respect to a conventional Booth radix-4 multiplier. Moreover, when combining it with partial carry-save adders it is possible to reduce 36 percent execution time on average for several benchmarks, achieving a 32.7 percent reduction in the Energy Delay Product at the same time. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik |
IEEE Trans. Computers | 1 |
| 2016 | A Distributed Clustered Architecture to Tackle Delay Variations in Datapath SynthesisabstractDue to the necessity of handling unexpected events in execution time, e.g., to support process variations, new mechanisms for dealing with every possible behavior of the datapath must be developed. Conventional centralized controllers can only handle very few dynamic events. Distributed controllers, on the other hand, are able to support every combination of events. These controllers are composed of several finite state machines, which are interconnected via a global coordinator. The use of this type of controller obliges to check the hazards between operations in run time, which entails some penalty in the controller complexity. In this paper, a new methodology for deploying a distributed controller over a set of clusters is presented. A register binding algorithm specially suited for distributed controllers has also been developed. It combines a clustering method and a least recently used policy to reduce the number of hazards in run time. Furthermore, our methodology allows the exploration of different solutions by tuning the input parameters of the binding algorithm. Several studies evaluating the execution time and area tradeoffs are presented to support our techniques. Results show that for some cases it is possible to reduce more than 50% the expected execution time, at the expense of a slight area increase. Alberto A. Del Barrio, Jason Cong, Román Hermida |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Ultra-low-power adder stage design for exascale floating point unitsabstractCurrently, the most powerful supercomputers can provide tens of petaflops. Future many-core systems are estimated to provide an exaflop. However, the power budget limitation makes these machines still unfeasible and unaffordable. Floating Point Units (FPUs) are critical from both the power consumption and performance points of view of today's microprocessors and supercomputers. Literature offers very different designs. Some of them are focused on increasing performance no matter the penalty, and others on decreasing power at the expense of lower performance. In this article, we propose a novel approach for reducing the power of the FPU without degrading the rest of parameters. Concretely, this power reduction is also accompanied by an area reduction and a performance improvement. Hence, an overall energy gain will be produced. According to our experiments, our proposed unit consumes 17.5%, 23% and 16.5% less energy for single, double and quadruple precision, with an additional 15%, 21.5% and 14.5% delay reduction, respectively. Furthermore, area is also diminished by 4%, 4.5 and 5%. Alberto A. Del Barrio, Nader Bagherzadeh, Román Hermida |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Multispeculative additive trees in high-level synthesisabstractMultispeculative Functional Units (MSFUs) are arithmetic functional units that operate using several predictors for the carry signal. The carry prediction helps to shorten the critical path of the functional unit. The average performance of these units is determined by the hit rate of the prediction. In spite of utilizing more than one predictor, none or only one additional cycle is enough for producing the correct result in the majority of the cases. In this paper we present multispeculation as a way of increasing the performance of tree structures with a negligible area penalty. By judiciously introducing these structures into computation trees, it will only be necessary to predict in certain selected nodes, thus minimizing the number of operations that can potentially mispredict. Hence, the average latency will be diminished and thus performance will be increased. Our experiments show that it is possible to improve on average 24% and 38% execution time, when considering logarithmic and linear modules, respectively. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik, Jose Manuel Mendias, María C. Molina |
DATE | 1 |
| 2013 | Exploring the energy efficiency of Multispeculative AddersabstractVariable Latency Adders are attracting strong interest for increasing performance at a low cost. However, most of the literature is focused on achieving a good area-delay tradeoff. In this paper we consider multispeculation as an alternative for designing adders with low energy consumption, while offering better performance than the corresponding non-speculative ones. Instead of introducing more logic to accelerate the computation, the adder is split into several fragments which operate in parallel, and whose carry-in signals are provided by predictor units. On the one hand, the critical path of the module is shortened, and on the other hand the frequent useless glitches produced in the carry propagation structure are diminished. Hence, this will be translated into an overall energy reduction. Several experiments have been performed with linear and logarithmic adders, and results show energy savings by up to 90% and 70%, respectively, while achieving an additional execution time decrease. Furthermore, when utilized in whole datapaths with current control techniques, it is possible to reduce execution time by 24.5% (34% best case) and energy by 32% (48% best case) on average. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik |
ICCD | 1 |
| 2013 | A fragmentation aware High-Level Synthesis flow for low power heterogenous datapaths
Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
Integr. | 1 |
| 2012 | Multispeculative Addition Applied to Datapath SynthesisabstractAddition is the key arithmetic operation in most digital circuits and processors. Therefore, their performance and other parameters, such as area and power consumption, are highly dependent on the adders' features. In this paper, we present multispeculation as a way of increasing adders' performance with a low area penalty. In our proposed design, dividing an adder into several fragments and predicting the carry-in of each fragment enables computing every addition in two very short cycles at the most, with 99% or higher probability. Furthermore, based on multispeculation principles, we propose a new strategy for implementing addition chains and hiding most of the penalty cycles due to mispredictions, while keeping at the same time the resource sharing capabilities that are sought in high-level synthesis. Our results show that it is possible to build linear and logarithmic adders more than$4.7\times$and$1.7\times$faster than the nonspeculative case, respectively. Moreover, this is achieved with a low area penalty (38% for linear adders) or even an area reduction (${-}8\%$for logarithmic adders). Finally, applying multispeculation principles to signal processing benchmarks that use addition chains will result in 25% execution time reduction, with an additional 3% decrease in datapath area with respect to implementations with logarithmic fast adders. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik, Jose Manuel Mendias, María C. Molina |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Power optimization in heterogenous datapathsabstractHeterogenous datapaths maximize the utilization of functional units (FUs) by customizing their widths individually through fragmentation of wide operands. In comparison, slices in large functional units in a homogenous datapath could be spending many cycles not performing actual useful work. Various fragmentation techniques demonstrated benefits in minimizing the total functional unit area. Upon a closer look at fragmentation techniques, we observe that the area savings achieved by heterogenous datapaths can be traded-off for power optimization. Our specific approach is to introduce choices for functional units with power/area trade-offs for different fragmentation and allocation choices, for reducing power consumption while satisfying the area constraint imposed on the heterogenous datapath. As low power FUs in literature produce an area penalty, a methodology must be developed in order to introduce them in the HLS flow while complying with the area constraint. We propose an allocation and module selection algorithms that pursue a trade-off between area and power consumption for fragmented datapaths under a total area constraint. Results show that it is possible to reduce power by 37% on average (49% in the best case). Moreover latency and cycle time will be equal or nearly the same as in the baseline case, which will lead to an energy reduction, too. Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
DATE | 1 |
| 2011 | A Distributed Controller for Managing Speculative Functional Units in High Level SynthesisabstractSpeculative functional units (SFUs) are arithmetic functional units that operate using a predictor for the carry signal. The carry prediction helps to shorten the critical path of the functional unit. The average case performance of these units is determined by the hit rate of the prediction. In case of mispredictions, the SFUs need to be coordinated by the datapath control mechanism to perform corrections and to maintain the datapath in the correct state. Devising a control mechanism for correcting mispredictions without adversely impacting overall performance is the most important challenge. In this paper, we present techniques for designing a datapath controller for seamless deployment of SFUs in high level synthesis. We have developed two techniques based on two main control paradigms: centralized and distributed control. The centralized approach stops the execution of the entire datapath for each misprediction and resumes execution once the correct value of the carry is known. The distributed approach decouples the functional unit suffering from the misprediction from the rest of the datapath. Hence, it allows the remainder of the functional units to carry on execution and be at different scheduling states at different times. We tested datapaths utilizing both linear structures and logarithmic structures for speculative arithmetic functional units. Our results show that it is possible to reduce execution time by as much as 38% (33% on average) for linear structures and by as much as 37.2% (25% on average) for logarithmic structures. Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | Using Speculative Functional Units in high level synthesisabstractSpeculative Functional Units (SFUs) enable a new execution paradigm for High Level Synthesis (HLS). SFUs are arithmetic functional units that operate using a predictor for the carry signal, which reduces the critical path delay. The performance of these units is determined by the success in the prediction of the carry value, i.e. the hit rate of the prediction. Hence SFUs reduce critical path at a low cost, but they cannot be used in HLS with the current techniques. In order to use them, it is necessary to include hardware support to recover from mispredictions of the carry signals. In this paper, we present techniques for designing a datapath controller for seamless deployment of SFUs in HLS. We have developed two techniques for this goal. The first approach stops the execution of the entire datapath for each misprediction and resumes execution once the correct value of the carry is known. The second approach decouples the functional unit suffering from the misprediction from the rest of the datapath. Hence, it allows the rest of the SFUs to carry on execution and be at different scheduling states at different times. Experiments show that it is possible to reduce execution time by as much as 38% and by 33% on average. Alberto A. Del Barrio, María C. Molina, Jose Manuel Mendias, Román Hermida, Seda Ogrenci Memik |
DATE | 1 |
| 2008 | Restricted Chaining and Fragmentation Techniques in Power Aware High Level SynthesisabstractA complete power-aware high-level synthesis algorithm is presented. It performs the schedule, resource allocation and binding of behavioral specifications. It overcomes the limitations of low-power algorithms and based on a bit-level timing model and a study of the target technology, tries to chain in the same cycle as many operations as possible. It also fragments the functional units, not the operations, for diminishing the required hardware. We also keep a minimum performance by estimating the cycle time while we are chaining operations. This way we obtain a reduction for both the static power and the dynamic one. We achieve an additional dynamic power reduction by studying the Hamming distance and applying partial or total commutative property. Experimental results on real circuits show great improvements in both power and energy consumption and performance over conventional low power algorithms. Alberto A. Del Barrio, María C. Molina, Jose Manuel Mendias, Esther Andres Perez, Román Hermida |
DSD | 1 |
| 2008 | Applying speculation techniques to implement functional unitsabstractThis paper justifies the use of estimation and prediction of carries to increase the performance of functional units built with the replication of full adders while keeping a low area penalization. Adders and multipliers are the most representative modules in this group of functional units. The use of these design techniques allows the implementation of modules with performance improvements ranging from 20% to 50% with only an area overheads around 5%. These functional units are suitable for asynchronous circuits but they could also be introduced in synchronous circuits with speculative techniques. The basic idea consists in estimating the carry out from some parts of the functional units, allowing every part to operate independently and in parallel. These modules are connected to build bigger ones. Results from simulations show that for some applications it is possible to make predictions even more accurate that the bit-based estimation. Predictions have also the advantage they can be introduced in the multipliers design, whether estimators cannot. These predictions are similar to the ones used in the branch prediction in a processor. Alberto A. Del Barrio, María C. Molina, Jose Manuel Mendias, Esther Andres Perez, Román Hermida, Francisco Tirado |
ICCD | 1 |