EDBT 2026 Demo / reviewers in the wild / expert
Prattay Chowdhury
dblp:241/0563
· DBLP profile ↗
9ranked-venue papers
7as first author
8since 2021 · last 2023
0000-0003-1018-7836ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Approximating HW Accelerators through Partial Extractions onto Shared Artificial Neural NetworksabstractOne approach that has been suggested to further reduce the energy consumption of heterogenous Systems-on-Chip (SoCs) is approximate computing. In approximate computing the error at the output is relaxed in order to simplify the hardware and thus, achieve lower power. Fortunately, most of the hardware accelerators in these SoCs are also amenable to approximate computing. Prattay Chowdhury, Jorge Castro-Godínez, Benjamin Carrión Schäfer |
ASP-DAC | 1 |
| 2023 | Application Specific Approximate Behavioral ProcessorabstractMany applications require simple controllers that continuously run the same application. These applications are often found in battery operated embedded systems that require to be ultra-low power (ULP) and are very price sensitive. Some examples include IoT devices of different nature and medical devices. Currently, these systems rely on off-the-shelf general-purpose microprocessors. One of the problems of using these processors, is that not all of the resources are needed for a specific application. Furthermore, because of the regularity of the workloads running on these systems there is a large opportunity to optimize the processor by pruning those unused resources to achieve lower area (cost) and power. Moreover, these processors can be specified at the behavioral level and use High-Level Synthesis (HLS) to generate an efficient Register Transfer Level (RTL) description. This opens a window to additional optimizations as the processor implementation is fully re-optimized during the HLS process. Also, many applications running on these embedded systems tolerate imprecise outputs. These include image processing and digital signal processing (DSP) applications. This opens the door to further optimizations in the context of approximate computing. To address these issues, this work presents a methodology to customize a behavioral RISC processor automatically for a given workload such that its area and power are significantly reduced as compared to the original, general-purpose processor. First, generating a bespoke processor that leads to the exact output as compared to the original general-purpose one and then by approximating it allowing a certain level of error at the output. Compared to previous work that customizes a given processor at the gate netlist only, our proposed method shows significant benefits. In particular, this work shows that raising the level of abstraction reduces the area and power by 78.3% and 70.1% for the exact solution on average, and further reduces the area by an additional 10.0% and 16.5% for the approximate version tolerating a maximum of 10% and 20% output errors respectively. Qilin Si, Prattay Chowdhury, Rohit Sreekumar, Benjamin Carrión Schäfer |
IEEE Trans. Sustain. Comput. | 2 |
| 2022 | Predictive Model Attack for Embedded FPGA Logic LockingabstractWith most VLSI design companies now being fabless it is imperative to develop methods to protect their Intellectual Property (IP). One approach that has become very popular due to its relative simplicity and practicality is logic locking. One of the problems with traditional locking mechanisms is that the locking circuitry is built into the netlist that the VLSI design company delivers to the foundry which has now access to the entire design including the locking mechanism. This implies that they could potentially tamper with this circuitry or reverse engineer it to obtain the locking key. One relatively new approach that has been coined logic locking through omission, or hardware redaction, maps a portion of the design to an embedded FPGA (eFPGA). The bitstream of the eFPGA now acts as the locking key. This new approach has been shown to be more secure as the foundry has no access to the bitstream during the manufacturing stage. The obvious drawbacks are the increase in design complexity and the area and performance overheads associated with the eFPGA. In this work we propose, to the best of our knowledge, the first attack on these type of new locking mechanisms by substituting the exact logic mapped onto the eFPGA by a synthesizable predictive model that replicates the behavior of the exact logic. We show that this approach is applicable in the context of approximate computing where hardware accelerators tolerate certain degree of errors at their outputs. Experimental results show that our proposed approach is very effective finding suitable predictive models while simultaneously reducing the overall power consumption. Prattay Chowdhury, Chaitali Sathe, Benjamin Carrión Schäfer |
ISLPED | 1 |
| 2022 | Leveraging Automatic High-Level Synthesis Resource Sharing to Maximize Dynamical Voltage Overscaling with Error ControlabstractApproximate Computing has emerged as an alternative way to further reduce the power consumption of integrated circuits (ICs) by trading off errors at the output with simpler, more efficient logic. So far the main approaches in approximate computing have been to simplify the hardware circuit by pruning the circuit until the maximum error threshold is met. One of the critical issues, though, is the training data used to prune the circuit. The output error can significantly exceed the maximum error if the final workload does not match the training data. Thus, most previous work typically assumes that training data matches with the workload data distribution. In this work, we present a method that dynamically overscales the supply voltage based on different workload distribution at runtime. This allows to adaptively select the supply voltage that leads to the largest power savings while ensuring that the error will never exceed the maximum error threshold. This approach also allows restoring of the original error-free circuit if no matching workload distribution is found. The proposed method also leverages the ability of High-Level Synthesis (HLS) to automatically generate circuits with different properties by setting different synthesis constraints to maximize the available timing slack and, hence, maximize the power savings. Experimental results show that our proposed method works very well, saving on average 47.08% of power as compared to the exact output circuit and 20.25% more than a traditional approximation method. Prattay Chowdhury, Benjamin Carrión Schäfer |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | Unlocking Approximations through Selective Source Code TransformationsabstractApproximate computing has emerged as a powerful alternative to further reduce the power of integrated circuits (ICs). The main idea is to simplify the software or hardware circuit to trade off lower power with a certain output error. This work presents a method to unlock approximations that under traditional approaches would have been ignored. In particular, we focus on Variable to Variable (V2V) and Variable to Constant (V2C) approximations in behavioral descriptions for either embedded software or High-Level Synthesis (HLS). The main idea in V2V and V2C is to substitute the computation of a variable by either another variable computed previously in the code or a constant to trade-off code size or circuit size with output error. Although these approximations are very powerful, we observed that they rarely occur. To enable these powerful approximations, we propose an automatic source code refactoring method combined with the selective substitution of portions of the code with predictive models. To maximize the savings in terms of code size, area, or power, we also propose an automated search method that finds the smallest possible predictive model for a given error threshold. Experimental results targeting a TI MSP430 processor as well as HLS show that our proposed method is very effective. Prattay Chowdhury, Benjamin Carrión Schäfer |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | Special Session: ADAPT: ANN-ControlleD System-Level Runtime Adaptable APproximate CompuTingabstractApproximate computing has shown to be an effective approach to generate smaller and more power-efficient circuits by trading the accuracy of the circuit vs. area and/or power. So far, most work on approximate computing has focused on specific components within a system. It severely limits the approximation potential as most Integrated Circuits (ICs) are now complex heterogeneous systems. One additional limitation of current work in this domain is they assume that the training data matches the actual workload. This is nevertheless not always true as these complex Systems-on-Chip (SoCs) are used for a variety of different applications. To address these issues, this work investigates if lower-power designs can be found through mixing approximations across the different components in the SoC as opposed to only aggressively approximating a single component. The main hypothesis is that some approximations amplify across the system, while others tend to cancel each other out, thus, allowing to maximize the power savings while meeting the given maximum error threshold. To investigate this, we propose a method called ADAPT. ADAPT uses a neural network-based controller to dynamically adjust the supply voltage (Vdd) of different components in SoC at runtime based on the actual workload. Prattay Chowdhury, Benjamin Carrión Schäfer |
ICCD | 1 |
| 2021 | BEACON: BEst Approximations for Complete BehaviOral HeterogeNeous SoCsabstractApproximate computing has shown to be an effective approach to generate smaller and more power-efficient circuits by trading the accuracy of the circuit vs. area/power. So far, most work on approximate computing has focused on specific components within a system. This severely limits the approximation potential as most Integrated Circuits (ICs) are now complex heterogeneous systems. This paper investigates if lower-power designs can be found through mixing approximations across the different components in the SoC as opposed to only aggressively approximating a single component. The main hypothesis is that some approximations amplify across the system, while others tend to cancel each other out, thus, allowing to maximize the power savings while meeting the given maximum error threshold. In this work, we consider the Analog-to-Digital Converter (ADC), CPU, hardware accelerators, and interconnect between all these components. Moreover, to quickly measure the effect of different approximation mixes, we have developed a framework that allows generating complete SoCs at the behavioral level through a bus generator and a library of synthesizable bus interfaces. This enables the use of fast simulation models (transaction and cycle-accurate) to accurately measure the error at the system’s output while measuring the benefit in terms of area or energy reduction of different mixes of approximations. Experimental results show that taking into account the entire system as oppose to only individual components leads to an additional average energy savings of 14% to 17% for different maximum error thresholds and the best case up to 39%. Prattay Chowdhury, Benjamin Carrión Schäfer |
ISLPED | 1 |
| 2021 | Estimating Operational Age of an Integrated Circuit
Prattay Chowdhury, Ujjwal Guin, Adit D. Singh, Vishwani D. Agrawal |
J. Electron. Test. | 1 |
| 2020 | Bespoke Behavioral ProcessorsabstractMany emerging applications require simple controllers that run the exact same application continuously. These include medical devices and IoTs of different nature. Because of the nature of these applications, they have to be ultra-low power and small. Most of the applications mapped onto low-power processors are computationally inexpensive, thus, amenable to be executed on a simple microprocessor. One of the problems of using a general purpose processor, is that not all of the resources are needed for a specific application, thus, there is a large potential for simplifying the processor to achieve lower area and power (static and dynamic). In addition, these processors can be specified at the behavioral level and use High-Level Synthesis (HLS) to generate the RTL automatically. This opens a window to additional optimizations as the processor can be pruned and re-synthesized at different VLSI design levels in order to obtain a smaller and more power-efficient processor. This work presents a methodology to customize a behavioral RISC processor automatically for a given workload such that its area and power are significantly reduced as compared to the original processor. Compared to previous work that customizes a given processor at the gate netlist only, our proposed method also shows significant benefits. In particular, we show that raising the level of abstraction reduces the area and power by 78.3 % and 70.1%. Rohit Sreekumar, Prattay Chowdhury, Benjamin Carrión Schäfer |
ICCD | 2 |