Shikha Goel

dblp:24/8395 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0000-8744-1789ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 EXPRESS: A Framework for Execution Time Prediction of Concurrent CNNs on Xilinx DPU Accelerator
abstract
Deep learning Processor Unit (DPU) is a highly configurable CNN accelerator that supports a variety of CNNs and can be implemented with multiple instances on the same FPGA. Many applications deploy concurrent execution of different CNNs and in such a setting, an execution time predictor can help “optimize” the DPU configurations to meet the performance requirements of different tasks. We characterize CNN execution on DPUs and reduce the variability in execution time due to interference from the operating system. Subsequently, we propose a machine learning-based framework (EXPRESS) to predict the execution time of any given CNN on a DPU configuration, considering CNN, DPU, and bus characteristics. We improvise EXPRESS to support heterogeneous CNNs in EXPRESS-2.0 by making features independent of the number of CNNs. Our entire experimentation is based on data from a real FPGA board for 16 standard CNNs. Our frameworks, EXPRESS and EXPRESS-2.0, significantly outperform state-of-the-art by achieving an average execution time prediction error of 2.2% and 0.7%, respectively. We illustrate the effectiveness of this low prediction error for design space exploration, which is very useful for embedded system application developers.
Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan
ACM Trans. Embed. Comput. Syst.1
2022 EXPRESS: CNN EXecution Time PREdiction for DPU DeSign Space Exploration
abstract
Deep learning Processor Units (DPUs) from Xilinx are design-time configurable CNN accelerators for FPGAs. We propose EXPRESS, which predicts the execution time of any given CNN on a DPU. EXPRESS incorporates the effect of bus connections into prediction. As a DPU is invoked by a host CPU to process a CNN layer by layer, EXPRESS considers the CPU and the DPU execution time for predicting the end-to-end processing time. EXPRESS has an average prediction error of 2.2% and significantly outperforms state-of-the-art.
Shikha Goel, Rajesh Kedia, Rijurekha Sen, M. Balakrishnan
FPT1
2021 EnergyNN: Energy Estimation for Neural Network Inference Tasks on DPU
abstract
Convolutional Neural Networks (CNNs) are increasingly becoming popular in embedded and energy limited mobile applications. Hardware designers have proposed various accelerators to speed up the execution of CNNs on embedded platforms. Deep Learning Processor Unit (DPU) is one such generic CNN accelerator for Xilinx platforms that can execute any CNN on one or more DPUs configured on an FPGA. In a period of rapid growth in CNN algorithms and the availability of multiple configurations of CNN accelerators (like DPU), the design space is expanding fast. These design points show significant trade-off in execution time, energy consumption and application performance measured in terms of accuracy. To be able to perform this trade-off, we propose a methodology for energy estimation of a CNN running on a DPU. We build an energy model using characteristics of few CNNs and use this model for energy prediction of other unseen CNNs. We evaluate our approach using 16 different standard and popular CNNs with an average prediction error of 9.9%. Energy estimation can be useful in various scheduling applications where one can choose from multiple CNNs based on its energy consumption. We demonstrate the utility of our approach in a drone that is deployed for detecting objects on the ground.
Shikha Goel, M. Balakrishnan, Rijurekha Sen
FPL1