Abhinav Goel

dblp:124/2914 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Impact of economic and socio-political risk factors on sovereign credit ratings
Abhinav Goel
Inf. Process. Manag.1
2024 Challenges and practices of deep learning model reengineering: A case study on computer vision
abstract
Abstract Context Many engineering organizations are reimplementing and extending deep neural networks from the research community. We describe this process as deep learning model reengineering. Deep learning model reengineering — reusing, replicating, adapting, and enhancing state-of-the-art deep learning approaches — is challenging for reasons including under-documented reference models, changing requirements, and the cost of implementation and testing. Objective Prior work has characterized the challenges of deep learning model development, but as yet we know little about the deep learning model reengineering process and its common challenges. Prior work has examined DL systems from a “product” view, examining defects from projects regardless of the engineers’ purpose. Our study is focused on reengineering activities from a “process” view, and focuses on engineers specifically engaged in the reengineering process. Method Our goal is to understand the characteristics and challenges of deep learning model reengineering. We conducted a mixed-methods case study of this phenomenon, focusing on the context of computer vision. Our results draw from two data sources: defects reported in open-source reeengineering projects, and interviews conducted with practitioners and the leaders of a reengineering team. From the defect data source, we analyzed 348 defects from 27 open-source deep learning projects. Meanwhile, our reengineering team replicated 7 deep learning models over two years; we interviewed 2 open-source contributors, 4 practitioners, and 6 reengineering team leaders to understand their experiences. Results Our results describe how deep learning-based computer vision techniques are reengineered, quantitatively analyze the distribution of defects in this process, and qualitatively discuss challenges and practices. We found that most defects (58%) are reported by re-users, and that reproducibility-related defects tend to be discovered during training (68% of them are). Our analysis shows that most environment defects (88%) are interface defects, and most environment defects (46%) are caused by API defects. We found that training defects have diverse symptoms and root causes. We identified four main challenges in the DL reengineering process: model operationalization, performance debugging, portability of DL operations, and customized data pipeline. Integrating our quantitative and qualitative data, we propose a novel reengineering workflow. Conclusions Our findings inform several conclusion, including: standardizing model reengineering practices, developing validation tools to support model reengineering, automated support beyond manual model reengineering, and measuring additional unknown aspects of model reengineering.
Wenxin Jiang 0001, Vishnu Banna, Naveen Vivek, Abhinav Goel, Nicholas Synovic, George K. Thiruvathukal, James C. Davis 0001
Empir. Softw. Eng.4
2022 Efficient Computer Vision on Edge Devices with Pipeline-Parallel Hierarchical Neural Networks
abstract
Computer vision on low-power edge devices enables applications including search-and-rescue and security. State-of-the-art computer vision algorithms, such as Deep Neural Networks (DNNs), are too large for inference on low-power edge devices. To improve efficiency, some existing approaches parallelize DNN inference across multiple edge devices. How-ever, these techniques introduce significant communication and synchronization overheads or are unable to balance workloads across devices. This paper demonstrates that the hierarchical DNN architecture is well suited for parallel processing on multiple edge devices. We design a novel method that creates a parallel inference pipeline for computer vision problems that use hierarchical DNNs. The method balances loads across the collaborating devices and reduces communication costs to facilitate the processing of multiple video frames simultaneously with higher throughput. Our experiments consider a representative computer vision problem where image recognition is performed on each video frame, running on multiple Raspberry Pi 4Bs. With four collaborating low-power edge devices, our approach achieves 3.21× higher throughput, 68% less energy consumption per device per frame, and a 58% decrease in memory when compared with existing sinaledevice hierarchical DNNs.
Abhinav Goel, Caleb Tung, Xiao Hu 0004, George K. Thiruvathukal, James C. Davis 0001, Yung-Hsiang Lu
ASP-DAC1
2022 Directed Acyclic Graph-based Neural Networks for Tunable Low-Power Computer Vision
abstract
Processing visual data on mobile devices has many applications, e.g., emergency response and tracking. State-of-the-art computer vision techniques rely on large Deep Neural Networks (DNNs) that are usually too power-hungry to be deployed on resource-constrained edge devices. Many techniques improve DNN efficiency of DNNs by compromising accuracy. However, the accuracy and efficiency of these techniques cannot be adapted for diverse edge applications with different hardware constraints and accuracy requirements. This paper demonstrates that a recent, efficient tree-based DNN architecture, called the hierarchical DNN, can be converted into a Directed Acyclic Graph-based (DAG) architecture to provide tunable accuracy-efficiency tradeoff options. We propose a systematic method that identifies the connections that must be added to convert the tree to a DAG to improve accuracy. We conduct experiments on popular edge devices and show that increasing the connectivity of the DAG improves the accuracy to within 1% of the existing high accuracy techniques. Our approach requires 93% less memory, 43% less energy, and 49% fewer operations than the high accuracy techniques, thus providing more accuracy-efficiency configurations.
Abhinav Goel, Caleb Tung, Nicholas Eliopoulos, Xiao Hu 0004, George K. Thiruvathukal, James C. Davis 0001, Yung-Hsiang Lu
ISLPED1
2021 Dynamic Adjustment of Concurrent Neural Networks within Limited Power Thermal Constraints in Autonomous Driving
abstract
Autonomous driving vehicles need to perform many different tasks like lane detection, blind-spot detection, and cruise control. Vision perception systems in autonomous vehicles run many neural networks concurrently to perform these tasks and thus have heavy processing workloads. The processors in automotive systems have significant power and thermal constraints that are dynamically changing in various field conditions. Existing techniques adjust processor and memory frequencies to regulate workloads to meet the required constraints. When presented with tight constraints, these solutions reduce the hardware speed and consequently increase the inference time of all the neural networks. Slower neural network inference is a safety hazard because mission-critical tasks like collision prevention are performed less frequently. We define an index, QoP (Quality of Perception), to represent the quality of the system-level vision perception. QoP requires high accuracy and processing speed on mission-critical autonomous driving tasks. We present a structured framework to maximize the QoP of the vision perception system within different power constraints. Our framework dynamically reallocates computing resources to the vision tasks by adjusting hyper-parameters of neural networks in run-time. The proposed method improves the QoP by an order of magnitude, which is a significant benefit for the reliability and safety of autonomous driving. Our experiments show that, in case of 15% power budget decrease, the proposed framework successfully maintains normal vision operations with high QoP (only $\sim$2.5% loss) while conventional methods of operating frequency throttling fail to function due to massive QoP degradation ($\sim$95% loss).
Hee Jun Park, Abhinav Goel
ICMLA2
2021 Low-Power Multi-Camera Object Re-Identification using Hierarchical Neural Networks
abstract
Low-power computer vision on embedded devices has many applications. This paper describes a low-power technique for the object re-identification (reID) problem: matching a query image against a gallery of previously-seen images. State-of-the-art techniques rely on large, computationally-intensive Deep Neural Networks (DNNs). We propose a novel hierarchical DNN architecture that uses attribute labels in the training dataset to perform efficient object reID. At each node in the hierarchy, a small DNN identifies a different attribute of the query image. The small DNN at each leaf node is specialized to re-identify a subset of the gallery-only the images with the attributes identified along the path from the root to a leaf. Thus, a query image is re-identified accurately after processing with a few small DNNs. We compare our method with state-of-the-art object reID techniques. With a $\sim 4\%$ loss in accuracy, our approach realizes significant resource savings: 74% less memory, 72% fewer operations, and 67% lower query latency, yielding 65% less energy consumption.
Abhinav Goel, Caleb Tung, Xiao Hu 0004, James C. Davis 0001, George K. Thiruvathukal, Yung-Hsiang Lu
ISLPED1
2021 Modular Neural Networks for Low-Power Image Classification on Embedded Devices
abstract
Embedded devices are generally small, battery-powered computers with limited hardware resources. It is difficult to run deep neural networks (DNNs) on these devices, because DNNs perform millions of operations and consume significant amounts of energy. Prior research has shown that a considerable number of a DNN’s memory accesses and computation are redundant when performing tasks like image classification. To reduce this redundancy and thereby reduce the energy consumption of DNNs, we introduce the Modular Neural Network Tree architecture. Instead of using one large DNN for the classifier, this architecture uses multiple smaller DNNs (called modules ) to progressively classify images into groups of categories based on a novel visual similarity metric. Once a group of categories is selected by a module, another module then continues to distinguish among the similar categories within the selected group. This process is repeated over multiple modules until we are left with a single category. The computation needed to distinguish dissimilar groups is avoided, thus reducing redundant operations, memory accesses, and energy. Experimental results using several image datasets reveal the effectiveness of our proposed solution to reduce memory requirements by 50% to 99%, inference time by 55% to 95%, energy consumption by 52% to 94%, and the number of operations by 15% to 99% when compared with existing DNN architectures, running on two different embedded systems: Raspberry Pi 3 and Raspberry Pi Zero.
Abhinav Goel, Sarah Aghajanzadeh, Caleb Tung, Shuo-Han Chen, George K. Thiruvathukal, Yung-Hsiang Lu
ACM Trans. Design Autom. Electr. Syst.1
2020 Camera Placement Meeting Restrictions of Computer Vision
abstract
In the blooming era of smart edge devices, surveillance cameras have been deployed in many locations. Surveillance cameras are most useful when they are spaced out to maximize coverage of an area. However, deciding where to place cameras is an NP-hard problem and researchers have proposed heuristic solutions. Existing work does not consider a significant restriction of computer vision: in order to track a moving object, the object must occupy enough pixels. The number of pixels depends on many factors (How far away is the object? What is the camera resolution? What is the focal length?). In this study, we propose a camera placement method that identifies effective camera placement in arbitrary spaces and can account for different camera types as well. Our strategy represents spaces as polygons, then uses a greedy algorithm to partition the polygons and determine the cameras' locations to provide the desired coverage. Our solution also makes it possible to perform object tracking via overlapping camera placement. Our method is evaluated against complex shapes and real-world museum floor plans, achieving up to 85% coverage and 25% overlap.
Sarah Aghajanzadeh, Roopasree Naidu, Shuo-Han Chen, Caleb Tung, Abhinav Goel, Yung-Hsiang Lu, George K. Thiruvathukal
ICIP5
2020 Low-power object counting with hierarchical neural networks
abstract
Deep Neural Networks (DNNs) achieve state-of-the-art accuracy in many computer vision tasks, such as object counting. Object counting takes two inputs: an image and an object query and reports the number of occurrences of the queried object. To achieve high accuracy, DNNs require billions of operations, making them difficult to deploy on resource-constrained, low-power devices. Prior work shows that a significant number of DNN operations are redundant and can be eliminated without affecting the accuracy. To reduce these redundancies, we propose a hierarchical DNN architecture for object counting. This architecture uses a Region Proposal Network (RPN) to propose regions-of-interest (RoIs) that may contain the queried objects. A hierarchical classifier then efficiently finds the RoIs that actually contain the queried objects. The hierarchy contains groups of visually similar object categories. Small DNNs at each node of the hierarchy classify between these groups. The RoIs are incrementally processed by the hierarchical classifier. If the object in an RoI is in the same group as the queried object, then the next DNN in the hierarchy processes the RoI further; otherwise, the RoI is discarded. By using a few small DNNs to process each image, this method reduces the memory requirement, inference time, energy consumption, and number of operations with negligible accuracy loss when compared with the existing techniques.
Abhinav Goel, Caleb Tung, Sarah Aghajanzadeh, Isha Ghodgaonkar, Shreya Ghosh 0003, George K. Thiruvathukal, Yung-Hsiang Lu
ISLPED1
2018 CompactNet: High Accuracy Deep Neural Network Optimized for On-Chip Implementation
abstract
Recent research has focused on Deep Neural Networks (DNNs) implemented directly in hardware. However, larger DNNs require significant energy and area, thereby limiting their wide adoption. We propose a novel DNN quantization technique and a corresponding hardware solution, CompactNet that optimizes the use of hardware resources even further, through dynamic allocation of memory for each parameter. Experimental results for the MNIST and CIFAR-10 datasets, show that CompactNet reduces the memory requirement by over 80%, the energy requirement by 12-fold, and the area requirement by 7-fold, when compared to the conventional DNN. This is achieved with minimal degradation to the classification accuracy. We demonstrate that, CompactNet provides pareto-optimal designs to make trade-offs between accuracy and resource requirement. The applications of CompactNet can be extended to datasets like ImageNet, and into models like MobileNet.
Abhinav Goel, Zeye Liu 0001, R. D. (Shawn) Blanton
IEEE BigData1
2018 Robustness of the Counting Rule for Distributed Detection in Wireless Sensor Networks
abstract
We consider the problem of energy-efficient distributed detection to infer the presence of a target in a wireless sensor network and analyze its robustness to modeling uncertainties. The sensors make noisy observations of the target's signal power, which follows the isotropic power-attenuation model. Binary local decisions of the sensors are transmitted to a fusion center, where a global inference regarding the target's presence is made, based on the counting rule. We consider uncertain knowledge of: 1) the signal decay exponent of the wireless medium; 2) the power attenuation constant; and 3) the distance between the target and the sensors. For a given degree of uncertainty, we show that there exists a limit on the target's signal power below which the distributed detector fails to achieve the desired performance regardless of the number of sensors deployed. Simulation results are presented to determine the level of sensitivity of the detector to uncertainty in these parameters. The results throw light on the limits of robustness for distributed detection, akin to “SNR walls” for classical detection.
Abhinav Goel, Adarsh Patel, Kyatsandra G. Nagananda, Pramod K. Varshney
IEEE Signal Process. Lett.1