EDBT 2026 Demo / reviewers in the wild / expert
Jorge Castro-Godínez
dblp:218/1204
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4808-4904ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust DCNN: The impact of approximate multipliers in defending against adversarial attacks
Mohammad Javad Askarizadeh, Jorge Castro-Godínez, Ebrahim Farahmand, Ali Mahani 0001, Laura Cabrera Quiros, Carlos Salazar-García |
Future Gener. Comput. Syst. | 2 |
| 2023 | Approximating HW Accelerators through Partial Extractions onto Shared Artificial Neural NetworksabstractOne approach that has been suggested to further reduce the energy consumption of heterogenous Systems-on-Chip (SoCs) is approximate computing. In approximate computing the error at the output is relaxed in order to simplify the hardware and thus, achieve lower power. Fortunately, most of the hardware accelerators in these SoCs are also amenable to approximate computing. Prattay Chowdhury, Jorge Castro-Godínez, Benjamin Carrión Schäfer |
ASP-DAC | 2 |
| 2023 | Automatic Generation of Resource and Accuracy Configurable Processing ElementsabstractLow-power consumption and scarce computational resources limit the computation at the edge. Besides, the approximate computing paradigm reports promising techniques for designing accelerators to deal with inherent limitations of the edge, and high-level synthesis with C++ opens the opportunity to use meta-programming for specialisable generic design. This work proposes a framework for automatically generating synthesis-time configurable processing elements (PEs) for matrix multiplication-addition (GEMMA) and convolution. To evaluate our work, we perform a design exploration after varying data bit-width, operand sizes, and kernel sizes. Our analyses include resource consumption scaling, clocks-to-solution, design efficiency, and error distribution, presenting a comprehensive view of how the parameters affect the properties of our generic implementations. The GEMMA presented a trade-off between granularity vs efficiency , where large PEs with short data widths are favoured by the design efficiency, achieving, theoretically, up to 75 GMAC/s on a Xilinx XC7Z020 @ 100 MHz with an efficiency of 27%. For design efficiency, we propose a figure of merit to evaluate operations per second and resource utilisation with respect to the maximum achievable by the FPGA. Regarding the convolution PEs, we implemented two algorithms: a window-based spatial convolution and Winograd. The former is the best in terms of performance with 150 GMAC/s, reaching up to 47% of efficiency. Winograd also outperformed numerically using a 3× 3 kernel filter, presenting a mean error of 11.01% in 4-bits operands with a PSNR=16.28 dB, compared to the spatial convolution with 38.2% of mean error and PSNR=5.89 dB. Finally, we discuss how the error is mostly dependent on the PE’s parameters. In the GEMMA, the error depends on the matrix size, causing limitations in the PE scaling but still applicable to accelerators. The PEs developed during this research will lead to further granular approximate accelerator research. Luis G. León-Vega, Eduardo Salazar-Villalobos, Alejandro Rodriguez-Figueroa, Jorge Castro-Godínez |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | AxRSU: Approximate Radix-4 Squarer UnitabstractApproximate computing emerged as a design alternative to boost design efficiency by leveraging the intrinsic error resiliency of many applications. Several error-resilient and compute-intensive applications such as signal, image, and video processing, computer vision, and supervised machine learning perform mean squared error (MSE) estimation during the runtime demanding dedicated squarer logic units in their hardware accelerators. This work proposes an approximate Radix-4 squarer unit architecture (AxRSU). Our AxRSU proposal reduces the encoder complexity and the number of required partial products, which considerably boosts energy and circuit area savings. We demonstrate the AxRSU error-quality trade-off in an SSD (Sum Squared Difference) hardware accelerator as a case study targeting a video processing application. We offer a new Pareto front with eighth optimal AxRSU solutions ranging 52-97% of cross-correlation (i.e., accuracy) for savings of 15-47% in energy consumption and 12-32% in circuit area. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Jorge Castro-Godínez, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
ISCAS | 3 |
| 2021 | Multiple approximate instances in neural processing units for energy-efficient circuit synthesis: work-in-progressabstractWe present an architectural approach toward energy-efficient synthesis of circuits used in neural processing units. Neural network applications are shown to tolerate varying operand precisions between different inputs, accuracy targets, their phases, and learning methods, without significantly impacting the classification accuracy. Using multiple instances of systolic arrays at different precisions, we show that significant energy gains are possible beyond the conventional approach, using the same circuit for all precisions. Tanfer Alan, Jorge Castro-Godínez, Jörg Henkel |
CASES | 2 |
| 2021 | Automatic Floorplanning and Standalone Generation of Bitstream-Level IP CoresabstractPartially reconfigurable designs on field-programmable gate array (FPGA) bring an opportunity for developers to license third-party intellectual property (IP) cores. There are multiple IP licensing models that can be used by the FPGA IP market. Their focus is mainly on feasibility and security; however, two major challenges have been ignored by almost all of them. First, both academic or industrial tools do not provide a flow to generate IPs in a standalone environment. Second, these tools only offer manual floorplanning of the IPs, which is both time and performance inefficient. In this work, we present a framework, that can be used by multiple parties to generate different parts of a design independently, that are compatible with each other. It also provides automatic floorplanning based on mixed-integer linear programming (MILP) that considers the distribution of heterogeneous resources in modern FPGAs, with efficient resource utilization as the main objective. The proposed floorplanning is evaluated with benchmarks from the related work. Furthermore, a use case of internal and open-source designs is used for the validation and evaluation of the independent IP generation and the floorplanner. Nadir Khan, Jorge Castro-Godínez, Shixiang Xue, Jörg Henkel, Jürgen Becker 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Towards Quality-Driven Approximate Software Generation for Accurate Hardware: Work-in-ProgressabstractMany existing processor-based systems, especially using off-the-shelf components, cannot afford hardware modifications to embrace different approximate computing techniques proposed by the research community. In that case, it is mainly at the software level where error resiliency can be exploited efficiently. Although a multitude of approximate techniques can be applied at software level, they have been presented in isolation and little has been done to report their combined usage on different types of applications amenable to approximations. We present here AxSWGen, an automated quality-driven methodology to jointly explore and apply multiple approximate techniques to error-tolerant sections of applications. AxSWGen is implemented using LLVM compiler infrastructure. We present results of automated approximate software generated with AxSWGen and executed on a RISC-V processor (SiFive HiFive1 board), achieving up to 50% energy reduction for a 5% image degradation for an approximate Gaussian filter. Jorge Castro-Godínez, Muhammad Shafique 0001, Jörg Henkel |
CASES | 1 |
| 2020 | AxHLS: Design Space Exploration and High-Level Synthesis of Approximate Accelerators using Approximate Functional Units and Analytical ModelsabstractWith the emergence of approximate computing as a design paradigm, many approximate functional units have been proposed, particularly approximate adders and multipliers. These circuits compromise the accuracy of their results within a tolerable limit to reduce the required computational effort and energy requirements. However, for an ongoing number of such approximate circuits reported in the literature, selecting those that minimize the required resources for designing and generating an approximate accelerator from a high-level specification, while satisfying a defined accuracy constraint, is a joint high-level synthesis (HLS) and design space exploration (DSE) challenge. In this paper, we propose a novel automated framework for HLS of approximate accelerators using a given library of approximate functional units. Since repetitive circuit synthesis and gate-level simulations require a significant amount of time, to enable our framework, we present AxME, a set of analytical models for estimating the required computational resources when using approximate adders and multipliers in approximate designs. We propose DSEwam, a DSE methodology for error-tolerant applications, in which analytical models, such as AxME, are used to estimate resources needed and the accuracy of approximate designs. Furthermore, we integrate DSEwam into an HLS tool to automatically generate Pareto-optimal, or near Pareto-optimal, approximate accelerators from C language descriptions, for a given error threshold and minimization goal. We release our DSE framework as an open-source contribution, which will significantly boost the research and development in the field of automatic generation of approximate accelerators. Jorge Castro-Godínez, Julián Mateus-Vargas, Muhammad Shafique 0001, Jörg Henkel |
ICCAD | 1 |
| 2019 | ECAx: Balancing Error Correction Costs in Approximate AcceleratorsabstractApproximate computing has emerged as a design paradigm amenable to error-tolerant applications. It enables trading the quality of results for efficiency improvement in terms of delay, power, and energy consumption under user-provided tolerable quality degradation. Approximate accelerators have been proposed to expedite frequently executing code sections of error-resilient applications while meeting a defined quality level. However, these accelerators may produce unacceptable errors at run time if the input data changes or dynamic adjustments are made for a defined output quality constraint. State-of-the-art approaches in approximate computing address this issue by correctly re-computing those accelerator invocations that produce unacceptable errors; this is achieved by using the host processor or an alternate exact accelerator, which is activated on-demand. Nevertheless, such approaches can nullify the benefits of approximate computing, especially when input data variations are high at run time and errors due to approximations are above a tolerable threshold. As a robust and general solution to this problem, we propose ECAx, a novel methodology to explore low-overhead error correction in approximate accelerators by selectively correcting most significant errors, in terms of their magnitude, without losing the gains of approximations. We particularly consider the case of approximate accelerators built with approximate functional units such as approximate adders. Our novel methodology reduces the required exact re-computations on the host processor, achieving up to 20% performance gain compared to state-of-the-art approaches. Jorge Castro-Godínez, Muhammad Shafique 0001, Jörg Henkel |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | Compiler-driven error analysis for designing approximate acceleratorsabstractApproximate Computing has emerged as a design paradigm suitable to applications with inherent error resilience. This paradigm aims to reduce the associated computing costs (such as execution time, area, or energy) of exact calculations by reducing the quality of their results. Several approximate arithmetic circuits have been proposed, which can be used to implement hardware blocks such as approximate accelerators. However, to satisfy quality constraints in these accelerators, it is imperative to assess how the errors introduced by approximate circuits propagate through other exact and approximate computations, and finally accumulate at the output. This is, in particular, crucial to enable high-level synthesis of approximate accelerators. This work proposes a compiler-driven error analysis methodology to evaluate the behavior of errors generated from approximate adders in the design of approximate accelerators. We present CEDA, a tool to perform a static analysis of the error propagation. This tool uses #pragma-based annotated C/C++ source code as input. With these annotations, exact additions are replaced by approximate ones during the code analysis to estimate the error at the output. The error estimations produced by our tool are comparable to those obtained through simulations. Jorge Castro-Godínez, Sven Esser, Muhammad Shafique 0001, Santiago Pagani, Jörg Henkel |
DATE | 1 |