Ravikumar V. Chakaravarthy

dblp:282/4668 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2022 A Task Parallelism Runtime Solution for Deep Learning Applications using MPSoC on Edge Devices
abstract
AI on edge devices [1]–[4] are becoming increasing popular over the last few years. There are many research projects TVM [5] TensorFlow lite [6] that have focused on deployment and acceleration of AI/ML models on edge devices. These solutions have predominantly used data parallelism to accelerate AI/ML models on the edge device using operator fusion, nested parallelism, memory latency hiding [5] etc. to achieve best performance on the supported hardware backends. However, when the hardware supports multiple heterogenous hardware backends it becomes important to support task parallelism in addition to data parallelism to achieve optimal performance. Tasks level parallelism [7] [8] helps break down an AI/ML model into multiple tasks that can be scheduled across various heterogenous backends available in a multi-processor system on chip (MPSoC). In our proposed solution we take an AI/ML compute graph and break it into a directed acyclic graph (DAG) such that each node of the DAG represents a sub-graph of the original compute graph. The nodes of the DAG are generated using an auto-tuner to achieve optimal performance for the corresponding hardware backend. The nodes are compiled into a binary executable for the targeted hardware backend and we are extending our machine learning framework, XTA [9], to generate DAG. The XTA runtime will analyze the DAG and generate scheduling configuration. The nodes of the DAG are analyzed for dependencies and parallelized or pipelined accordingly. We are seeing a 30% improvement over the current solutions by parallelizing the execution of nodes in the DAG. The performance can be further optimized by using more hardware backend cores of the MPSoC to execute the nodes of the DAG in parallel, which is missing in the existing solutions.
Raghav Chakravarthy, Ravikumar V. Chakaravarthy
ASP-DAC3
2022 Auto-tuning of AI/ML Graphs for Optimal Performance in a Heterogenous Processor System
abstract
High-performance parallel computing is very important to ensure the efficient execution of deep learning algorithms. However, current research focuses more on parallel computing of homogeneous processors. Due to the differences in memory, communication units and computing capabilities of heterogeneous architectures, existing deep learning parallel computing architectures cannot achieve good performance improvement in heterogeneous environment. Current solutions use serial methods to execute algorithms on heterogeneous architectures, rely on developer’s development experience or use simple search strategies to perform tuning to achieve parallel computation. The existing solutions either cannot fully utilize the hardware resources of heterogeneous architectures to achieve performance optimization or cannot maximize performance due to the design of limited search space and low efficient search strategies.In our solution, we designed a new neural network automatic horizontal partitioning framework, which automatically splits the neural network horizontally into multiple sub-graphs based on a given target device or target device cluster, so that the neural network can be running sub-graphs in a pipelined or parallel manner on a target device or target device cluster. Our solution creates a larger search space compared to existing solutions and uses efficient exploration and cost evaluation methods.
Ravikumar V. Chakaravarthy, Raghav Chakravarthy, Siddharth Das
ICCD1
2021 Vision Control Unit in Fully Self Driving Vehicles using Xilinx MPSoC and Opensource Stack
abstract
Fully self-driving (FSD) vehicles are becoming increasing popular over the last few years and companies are investing significantly into its research and development. In the recent years, FSD technology innovators like Tesla, Google etc. have been working on proprietary autonomous driving stacks and have been able to successfully bring the vehicle to the roads. On the other end, organizations like Autoware Foundation and Baidu are fueling the growth of self-driving mobility using open source stacks. These organizations firmly believe in enabling autonomous driving technology for everyone and support developing software stacks through the open source community that is SoC vendor agnostic. In this proposed solution we describe a vision control unit for a fully self-driving vehicle developed on Xilinx MPSoC platform using open source software components.
Ravikumar V. Chakaravarthy, Hyun Kwon
ASP-DAC1
2020 Special Session: XTA: Open Source eXtensible, Scalable and Adaptable Tensor Architecture for AI Acceleration
abstract
Accelerator frameworks have gained prominence since the advent of AI applications. The limitation with current open source accelerator solutions is that it was not designed to be scalable and adaptable for commercial MPSoC products that have different network requirements and higher performance goals. We have implemented a new AI accelerator framework, XTA, derived from TVM-VTA which is a popular, first known, open source backend AI accelerator for Xilinx MPSoC. XTA is scalable and adaptable to various network types and workloads of AI applications. XTA is a multi-core architecture that can dynamically scale and adapt to a given AI problem at both hardware and software layers. At the hardware layer it can adapt to compute and memory configurations of the system and at the software layer it can hide hardware complexity and adapt to changing user workloads or data flows of a given AI problem. XTA also supports parallel, pipelined processing and autotuning of subgraphs in a MPSoC environment. We hope that with this Open Source AI accelerator, industry can not only push the performance limits but also quickly innovate new AI applications based on the flexibly the architecture provides. Simulator version of the XTA shows significant performance improvements over TVM-VTA for a wide range of networks and workloads.
Ravikumar V. Chakaravarthy
ICCD1